Agentic Compiler Engineer

Posted Yesterday
Be an Early Applicant
Paris, Île-de-France, FRA
Hybrid
Entry level
Gaming
The Role
Build AGCO, an agentic compiler for optimizing LLM execution across GPUs. Responsibilities include compiler and IR design, optimization passes, lowering, code generation, search-based implementation exploration, verification, GPU execution and profiling, LLM inference optimization, and measurement-driven optimization loops on real hardware. The role offers ownership across compilers, GPU systems, and inference, with the primary focus tailored to the candidate’s expertise.
Summary Generated by Built In
ABOUT KOG

Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels.

On the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming.

On the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs.

Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.

Our next major project is AGCO, our agentic compiler. AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference.

The team has 10 people, including 9 engineers and researchers and 4 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

 
WHAT YOU WILL WORK ON

You will work directly on AGCO.

The goal is to build a system that can explore ways to optimize LLM execution, generate changes, compile them, check correctness, run them on real hardware, measure the results, and use this feedback to guide the next optimization.

You will contribute to areas such as:

  • Compiler and IR design for representing and transforming LLM computations.

  • Optimization passes, lowering, and code generation.

  • Search methods for exploring different implementations and execution strategies.

  • Verification and correctness checks for generated changes.

  • GPU execution, profiling, and performance optimization.

  • LLM inference across operators, memory, parallelism, and communication.

  • Optimization loops that connect generated changes to measurements on real GPUs.

One direction we are exploring combines an IR, a verifier, a compiler, and a search optimizer. We plan to start with focused problems, build working prototypes, and extend the system from what we learn.

Your main area will depend on your experience, skills, and interests. You may focus more on compilers, GPU systems, or LLM inference while working closely with people across the full stack.

 
WHAT WE LOOK FOR

We look for engineers with deep technical expertise and original work in at least one area relevant to AGCO.

Relevant experience includes:

  • Compiler engineering, including optimization passes, IRs, lowering, code generation, LLVM, or MLIR.

  • GPU programming with CUDA, HIP, Metal, Vulkan, or similar technologies.

  • GPU performance work involving kernels, memory, synchronization, profiling, or hardware behavior.

  • LLM inference engines and performance optimization.

  • Attention, MoE, parallelism, communication, or other systems-level parts of LLM execution.

  • Formal verification, equivalence checking, SAT/SMT, or related methods.

  • Systems that generate, search, test, benchmark, or optimize code automatically.

We care about what you personally built and the technical decisions behind it. Strong candidates can explain the problem, their approach, the alternatives they explored, and how they measured the result.

We review technical work during the process. This can be public code, an upstream contribution, a paper, a thesis, a technical project, or a detailed write-up based on work you can share.

 
WHAT WE OFFER

You will join a small team building AGCO as a core part of Kog's technology.

  • Work at the intersection of compilers, GPU systems, and LLM inference.

  • Direct access to engineers working across the full inference stack.

  • A fast loop from an optimization idea to compilation, execution, verification, and measurement on real GPUs.

  • The opportunity to go deep in your strongest technical area while expanding into the other parts of the stack.

  • High ownership over technical decisions and systems that will shape how Kog optimizes LLM inference.

This role is based in Paris, and we are looking for candidates who can relocate to Paris and work closely with the team.

Skills Required

  • Deep technical expertise and original work in at least one area relevant to agentic compilers, GPU systems, or LLM inference
  • Experience with compiler engineering, GPU programming, GPU performance optimization, LLM inference, systems-level LLM execution, formal verification, or automated code optimization systems
  • Ability to explain technical problems, approaches, alternatives, and measured results
  • Public code, upstream contribution, paper, thesis, technical project, or detailed shareable technical write-up for review
  • Ability to relocate to Paris and work closely with the team
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
35 Employees
Year Founded: 1994

What We Do

KOG Studios is a South Korean video game developer based in Daegu that specializes in producing online free-to-play games, including Elsword, KurtzPel: Bringer of Chaos, and Grand Chase.

Similar Jobs

Formance Logo Formance

Solutions Engineer

Blockchain • Fintech • Payments • Software • Financial Services • Cryptocurrency
Hybrid
Paris, Île-de-France, FRA
34 Employees

Datadog Logo Datadog

Research Manager - Foundation & World Models

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Paris, Île-de-France, FRA
6500 Employees

monday.com Logo monday.com

Sales Engineer

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Paris, Île-de-France, FRA
3155 Employees

Adyen Logo Adyen

Project Operation Manager

Fintech • Payments • Financial Services
Easy Apply
Hybrid
Paris, Île-de-France, FRA
4771 Employees

Similar Companies Hiring

DraftKings Thumbnail
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Boston, MA
6400 Employees
bet365 Thumbnail
Digital Media • Gaming • Software • Esports • Automation
Denver, Colorado
10000 Employees
ARB Interactive Thumbnail
Gaming • Mobile • Software
Miami, Florida
190 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account