GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Posted Yesterday
Be an Early Applicant
5 Locations
Remote
Mid level
Artificial Intelligence • Edtech • Machine Learning • Professional Services
The Role
Reviews, debugs, evaluates, and optimizes GPU and accelerator kernels for AI workloads. Responsibilities include CUDA and Triton optimization, framework translation, hardware migration, operator fusion, profiling, benchmarking, numerical correctness validation, compilation and runtime debugging, and memory hierarchy optimization. The role assesses implementation quality, performance bottlenecks, hardware limitations, and technical task configuration while providing actionable feedback.
Summary Generated by Built In

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging

  • CUDA and Triton optimization

  • Translation between kernel frameworks

  • Hardware migration

  • Operator fusion

  • Performance profiling and benchmarking

  • Numerical correctness verification

  • Compilation and runtime debugging

  • Memory hierarchy optimization

  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For
  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

  • Strong experience with at least two of the following:

    • CUDA

    • Triton

    • NKI / AWS Neuron

    • Pallas / JAX

  • Strong understanding of GPU performance optimization

  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

  • Understanding of:

    • Memory bandwidth

    • Compute throughput

    • GPU occupancy

    • Shared memory

    • Register pressure

    • Memory coalescing

    • Bank conflicts

  • Strong understanding of floating-point numerical correctness and tolerance thresholds

  • Experience debugging kernel compilation and runtime issues

  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications

  • Translating kernels between CUDA, Triton, or other frameworks

  • Migrating kernels across hardware platforms

  • Debugging incorrect kernel implementations

  • Optimizing kernel performance

  • Fusing multiple operations into optimized kernels

Nice to Have
  • Experience across both NVIDIA GPU and custom accelerator ecosystems

  • Experience with AWS Trainium, TPU, JAX, or other accelerators

  • Compiler engineering experience

  • Familiarity with MLIR, XLA, or intermediate representation lowering

  • Contributions to GPU or ML kernel libraries

  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

  • Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For
  • Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations

  • Evaluating numerical tolerance thresholds

  • Reviewing kernel benchmarks and determining whether comparisons are fair

  • Identifying performance bottlenecks and optimization opportunities

  • Assessing whether performance targets are realistic given hardware limits

  • Reviewing kernel translations and hardware migrations

  • Identifying compilation, driver, memory, shape, and runtime issues

  • Determining whether technical tasks are genuinely difficult or incorrectly configured

  • Providing clear, actionable technical feedback

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Skills Required

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
  • Strong experience with at least two of CUDA, Triton, NKI/AWS Neuron, or Pallas/JAX
  • Strong understanding of GPU performance optimization
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
  • Understanding of memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts
  • Strong understanding of floating-point numerical correctness and tolerance thresholds
  • Experience debugging kernel compilation and runtime issues
  • Ability to distinguish software defects, environment problems, and genuine optimization challenges
  • Experience writing kernels from technical specifications, translating kernels, migrating across hardware, debugging incorrect implementations, optimizing performance, or fusing operations
  • Experience across NVIDIA GPU and custom accelerator ecosystems
  • Experience with AWS Trainium, TPU, JAX, or other accelerators
  • Compiler engineering experience
  • Familiarity with MLIR, XLA, or intermediate representation lowering
  • Contributions to GPU or ML kernel libraries
  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
  • Experience with AI model evaluation, RLHF, or technical benchmark development
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2022

What We Do

Anyone AI is an edtech startup dedicated to bridging the AI talent gap by investing in software developers from Latin America. The company provides intensive, hands-on training programs in Machine Learning and Artificial Intelligence, led by industry experts. By combining technical skill development with employability support, Anyone AI prepares professionals for global career opportunities, helping them transition into high-impact roles within the rapidly evolving AI and technology sectors.

Similar Jobs

InterSystems Logo InterSystems

Administrative Assistant

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Easy Apply
Remote
Chile
2100 Employees

Cloudflare Logo Cloudflare

Senior Customer Engineer, LATAM - MCR (Santiago, Chile).

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Chile
4400 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sr. Sales Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

Domino Data Lab Logo Domino Data Lab

Support Engineer

Artificial Intelligence • Machine Learning
Remote or Hybrid
10 Locations
200 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account