GPU Performance Engineer

Reposted 25 Days Ago
Be an Early Applicant
San Francisco, CA
In-Office
Senior level
Artificial Intelligence • Generative AI
The Role
Optimize GPU performance, debug issues, write custom kernels, and collaborate with ML engineers to enhance model serving efficiency.
Summary Generated by Built In

We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.

We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.
The Role
You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.
Key Responsibilities

  • Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation

  • Write high-performance CUDA and Triton kernels for critical model operations

  • Optimize cold start latency from seconds to milliseconds for our serving infrastructure

  • Tune memory access patterns, kernel fusion, and GPU utilization

  • Collaborate with ML engineers to optimize model implementations

  • Debug performance issues across the full stack from application to hardware

  • Implement custom memory pooling and allocation strategies

  • Share optimization techniques and build performance culture across teams

Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field

  • 5+ years systems programming experience with 3+ years focused on GPU optimization

  • Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)

  • Strong CUDA programming skills with production kernel development

  • Deep understanding of GPU architecture (memory hierarchy, SMs, warps)

  • Track record of achieving significant performance improvements (5-10x)

  • Experience with Python and C++ in production environments

We Value

  • Experience with Triton kernel development

  • Knowledge of CUTLASS or similar high-performance libraries

  • Background in ML-specific optimizations (attention, transformers)

  • RDMA/InfiniBand optimization experience

  • Contributions to GPU libraries or frameworks

  • Low-level debugging skills (PTX/SASS reading)

Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.

Top Skills

C++
Cuda
Nsight Systems
Nvprof
Python
Triton
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
San Francisco, CA
50 Employees

What We Do

Enabling the next billion AI video creators with Genmo

Similar Jobs

Anthropic Logo Anthropic

Performance Engineer - GPU

Artificial Intelligence • Natural Language Processing • Generative AI
In-Office
3 Locations
57 Employees
315K-560K Annually

NVIDIA Logo NVIDIA

GPU Performance Profiling Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office
2 Locations
21960 Employees
148K-288K Annually

Celonis Logo Celonis

Senior Consultant

Big Data • Information Technology • Productivity • Software • Analytics • Business Intelligence • Consulting
Hybrid
Redwood City, CA, USA
3000 Employees
119K-165K Annually

CrowdStrike Logo CrowdStrike

Counsel

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
3 Locations
10000 Employees
185K-255K Annually

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Standard Template Labs Thumbnail
Software • Information Technology • Artificial Intelligence
New York, NY
10 Employees
Scotch Thumbnail
Software • Retail • Payments • Fintech • eCommerce • Artificial Intelligence • Analytics
US
25 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account