Kog builds the Kog Inference Engine, a real-time inference engine for AI agents running on standard datacenter GPUs.
We co-design three layers: model architecture, inference engine, and low-level GPU kernels. We design the full inference stack around standard AMD and NVIDIA datacenter GPUs, from model architecture down to low-level kernels.
Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.
Our hot path removes framework and host overhead. On NVIDIA, we write CUDA and PTX by hand. On AMD, we use HIP and CDNA ISA.
The team has 11 people, including 10 engineers and researchers and 5 PhDs.
Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.
What you will work onThe GPU Engineer role focuses on low-level execution and GPU performance. Memory behavior, synchronization, latency, and hardware constraints shape the work.
You will work on:
Our monokernel pipeline, where the decode loop runs as one persistent GPU program from the first token to the last while the GPU keeps control of the hot path.
Low-level kernel optimization across AMD and NVIDIA, with both platforms treated as first-class targets.
Memory-bound execution paths, including the batch-size-1 GEMV regime as our primary target.
Profiling infrastructure that isolates bottlenecks inside a persistent GPU program and connects measurements to implementation decisions.
Inter-GPU communication through KCCL, our latency-focused communication layer.
Scaling the engine to third-party MoE models, with DeepSeek V4 as the current porting target.
Internal agents for GPU engineering and kernel optimization, built on the expertise and execution infrastructure developed by the GPU team.
We look for engineers who have worked below the framework layer on problems where hardware behavior or performance was central.
Relevant evidence includes:
CUDA, HIP, PTX, CDNA ISA, SASS, shaders, drivers, or comparable low-level systems work.
GPU kernels where you measured performance and can explain why a change moved the result.
Latency-sensitive, memory-bound, or synchronization-sensitive execution paths.
Profiling work that identified a real bottleneck and led to an implementation change.
Upstream contributions to inference engines, compilers, drivers, graphics systems, or other performance-critical projects.
Original technical work in graphics, game engines, video, Vulkan, drivers, HPC, or scientific computing at the hardware and performance layer.
PyTorch custom operations and Triton are relevant when the work shows hardware-level reasoning below the API layer.
We review a technical artifact during the process. This can be public code, a merged upstream contribution, a thesis, or a detailed technical write-up based on work you can share.
What we offerYou will join a small team at a stage where individual engineers can still shape the core technology, while the engine is advanced enough to start being tested against real customer workloads.
High individual impact, with direct ownership over technical decisions and systems that sit on the critical path of inference performance.
AMD and NVIDIA as first-class targets, giving you exposure to different GPU architectures, programming models, and hardware behaviors within the same inference engine.
A broad technical surface, where you can follow a performance problem from profiling and kernel execution through synchronization and inter-GPU communication.
A role at the foundation of our internal GPU engineering agents, where the expertise and systems built by the GPU team become the substrate for automated kernel optimization.
A Paris-based team with support for candidates outside the region: at least one week per month in Paris, with travel and accommodation covered by Kog.
Skills Required
- Proven experience writing GPU kernels where performance is the central constraint (show code).
- Able to provide code samples or public repositories demonstrating low-level GPU work.
- Willingness to spend at least 50% of time in the Paris office.
- Experience developing PyTorch custom ops for GPU (acceptable starting point).
- Experience with inline PTX or CDNA ISA in public repositories.
- Experience with latency-sensitive execution paths and low-batch inference optimizations (MBU vs MFU).
- Background with inference engine components (GEMM, attention kernels, collectives, decoding pipeline).
What We Do
KOG Studios is a South Korean video game developer based in Daegu that specializes in producing online free-to-play games, including Elsword, KurtzPel: Bringer of Chaos, and Grand Chase.







