KERNEL ENGINEER

Posted 2 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Mid level
Artificial Intelligence • Information Technology • Machine Learning • Software
The Role
Write and optimize GPU kernels for training and inference, profile workloads with hardware counters, co-design kernels with researchers, integrate and benchmark kernels in training/serving stacks, and maintain kernel quality while sharing expertise across the team.
Summary Generated by Built In

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You'll write and optimize the GPU kernels and supporting systems software that makes our training and inference workloads fast. This is deep, low-level work (performance counters, memory bandwidth, warp-level scheduling) applied to the specific shapes and patterns our models actually use.

We hire kernel engineers because the gap between "this works" and "this is fast on the hardware we have" is enormous, and that gap directly bounds what our researchers can try. You'll close that gap.

WHAT YOU'LL DO

  • Write and optimize GPU kernels (CUDA, ROCm, Triton, or similar) for training and inference workloads: attention variants, MoE layers, custom activations, communication primitives

  • Profile real workloads with hardware counters and translate findings into specific kernel-level optimizations

  • Co-design kernels with the research teams, when the kernel and the algorithm need to change together, you participate in both

  • Integrate optimized kernels into our training and serving stacks; benchmark before and after; verify the win is real end-to-end

  • Maintain kernel quality over time as hardware, frameworks, and workloads shift underneath

  • Spread kernel-level fluency across the team; we want this expertise shared, not siloed

WHAT WE'RE LOOKING FOR

  • 4+ years writing performant GPU kernels (CUDA, ROCm, Triton, or production-grade equivalent)

  • Hardware-level fluency: memory hierarchy, occupancy, register pressure, tensor cores, warp scheduling

  • Profiling fluency (Nsight, ncu, or comparable tools) and the discipline to measure before changing

  • Track record of shipping kernel-level optimizations that moved a measurable metric in a real system

  • Strong systems expertise: you understand how kernels live inside larger frameworks and how integration choices affect end-to-end performance

  • Comfortable reading framework-level Python and C++ around your kernels

NICE TO HAVE

  • Open-source contributions to kernel libraries, compilers, or ML frameworks

  • Experience with multiple accelerator architectures (different GPU families, TPUs, custom ASICs), preferably AMD GPUs

  • Familiarity with collective communication primitives (NCCL or equivalent)

  • Compiler or runtime background

THIS ROLE IS PROBABLY NOT FOR YOU IF

  • You haven't gotten your hands dirty at the kernel level: this isn't a higher-level systems role rebranded

  • You want to stay narrowly in one library; we expect breadth across the kernel surface our models actually use

  • Performance work without measurable end-to-end impact frustrates you

Skills Required

  • 4+ years writing performant GPU kernels (CUDA, ROCm, Triton, or equivalent)
  • Hardware-level fluency (memory hierarchy, occupancy, register pressure, tensor cores, warp scheduling)
  • Profiling fluency with tools such as Nsight or ncu and disciplined measurement practices
  • Proven track record of shipping kernel-level optimizations that moved measurable metrics in real systems
  • Strong systems expertise; understanding how kernels integrate with larger frameworks and affect end-to-end performance
  • Comfortable reading framework-level Python and C++ surrounding kernels
  • Open-source contributions to kernel libraries, compilers, or ML frameworks
  • Experience with multiple accelerator architectures (different GPU families, TPUs, custom ASICs), preferably AMD GPUs
  • Familiarity with collective communication primitives (NCCL or equivalent)
  • Compiler or runtime background
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
8 Employees

What We Do

MakerMaker.AI is an innovative company building autonomous research agents focused on recursive self-improvement. They develop sophisticated multi-agent systems that can independently propose, run, and analyze complex machine learning experiments. By creating agents that have the ability to build other agents, MakerMaker.AI seeks to accelerate the development of artificial intelligence and push the boundaries of autonomous research in machine learning.

Similar Jobs

CoreWeave Logo CoreWeave

Senior Software Engineer

Cloud • Information Technology • Machine Learning
In-Office
2 Locations
1450 Employees
182K-242K Annually

NVIDIA Logo NVIDIA

Senior Inference Engineer, GPU Kernel Optimization

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office
4 Locations
21960 Employees
184K-288K Annually

Etched Logo Etched

Artificial Intelligence Engineer

Artificial Intelligence • Hardware • Software
In-Office
San Jose, CA, USA
53 Employees
150K-225K Annually
In-Office
Walnut Creek, CA, USA
17787 Employees
130K-160K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account