Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, and low-level GPU kernels.
On the model side, we developed Laneformer 2B, our current coding model, and Delayed Tensor Parallelism (DTP), a Transformer architecture that delays communication so it overlaps with useful computation and weight streaming.
On the systems side, the Kog Inference Engine runs this co-designed stack on standard AMD and NVIDIA datacenter GPUs.
Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.
Our approach is to design the model and its execution together, so architecture, communication, memory behavior, and GPU execution can be optimized as parts of the same system.
The team has 11 people, including 10 engineers and researchers and 5 PhDs.
Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.
What you will work onThe Research Engineer role focuses on model architecture designed around inference. Communication, memory, parallelism, and execution constraints shape the research.
You will work on:
Inference-aware model architecture across attention, routing, residual structure, and parallelism.
Extensions to Laneformer 2B and DTP.
Architecture choices for large MoE models, including how routing, communication, and execution interact at inference time.
Post-training methods that change model architecture or inference behavior.
Model morphing, changing the size or structure of a trained model while reusing its learned weights.
Experiments that connect architectural hypotheses to measurable model and inference behavior.
Publishing research results and turning successful ideas into systems that run inside the Kog inference stack.
We look for researchers who have designed or materially changed model mechanisms and can connect architectural decisions to how models execute at inference time.
Relevant evidence includes:
Original work on attention, routing, residual connections, parallelism, or other Transformer mechanisms.
Model architecture work where communication, memory, hardware, or inference constraints shaped the design.
Research on Transformers or MoE models where you owned a meaningful part of the architecture or experimental direction.
Experiments built around a clear hypothesis, a decisive measurement, and an update to the model or research direction.
Post-training work that changed model structure or inference behavior.
Work on model scaling, model morphing, or techniques that reuse learned weights across architectural changes.
We review a technical artifact during the process. This can be a paper, public code, a thesis, a research project, or a detailed technical write-up based on work you can share.
You will join a small team at a stage where individual researchers can still shape the core technology, as our technology starts being used on real customer workloads.
High individual impact, with research decisions that can directly influence both model architecture and the systems built to execute it.
A short path from research to implementation, with architectural ideas evaluated through both model behavior and real inference performance.
The ability to take an architectural idea from hypothesis to experiment, implementation, and real inference behavior.
The opportunity to extend research such as Laneformer and DTP, explore new architectural directions, and publish results.
A Paris-based team with support for candidates outside the region: at least one week per month in Paris, with travel and accommodation covered by Kog.
Skills Required
- Demonstrable work on complex AI problems (paper, repository, thesis, or substantial side project)
- Experience adapting or modifying existing model architectures
- Fluency in Transformers and Mixture of Experts (MoE) architectures
- Understanding of communication structure, layer dependencies, and how they affect inference
- Experience scaling models for inference (routing, expert parallelism, communication patterns)
- Experience with post-training methods such as fine-tuning, quantization, or preference optimization
- Ability to write and submit research papers and present at conferences
What We Do
KOG Studios is a South Korean video game developer based in Daegu that specializes in producing online free-to-play games, including Elsword, KurtzPel: Bringer of Chaos, and Grand Chase.








