Deep Learning Performance Architect, CUTLASS DSL

Reposted 20 Days Ago
Be an Early Applicant
2 Locations
In-Office
Mid level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Design, develop, and optimize CUTLASS DSL for high-performance GPU kernels. Advance MLIR dialects and improve kernel compilation speed while ensuring performance. Collaborate across teams to implement optimizations.
Summary Generated by Built In

Are you passionate about programming languages, compiler technology, and GPU performance? Do you want to help shape the future of high-performance kernel development for AI? We are looking for outstanding engineers to build CUTLASS DSL, a Python-native language for GPU kernel development, along with the MLIR dialects and lowering passes behind it. In this role, you will also help accelerate kernel compilation while delivering performance comparable to CUTLASS C++, enabling efficient hardware-software co-design for NVIDIA's next generation of AI platforms. 

What you'll be doing: 

  • Design, develop, and optimize CUTLASS DSL, a Python-native language for high-performance GPU kernel development 

  • Build and advance the MLIR dialects, lowering passes, and code generation flows that power the CUTLASS DSL stack  

  • Drive innovations that improve kernel compilation speed while maintaining performance on par with CUTLASS C++ 

  • Collaborate closely with architecture, research, software product teams, and the open-source community to bring cutting-edge optimizations into real products 

What we need to see: 

  • MS, PhD, or equivalent experience in Computer Science, Software Engineering, or a related field 

  • 2+ years of relevant work experience 

  • Excellent programming skills in Python and strong proficiency in C++ 

  • Hands-on experience with DSLs, compilers, or code generation systems 

  • Strong command of the MLIR/LLVM stack, including IR design and pass optimization 

  • Strong communication skills and the ability to thrive in a highly collaborative environment 

Ways to stand out from the crowd: 

  • Deep understanding of the CUDA GPU programming model, GPU microarchitecture, and performance analysis and optimization techniques 

  • Familiarity with key high-performance computing abstractions such as Layout, Tile, MMA, and TMA in the CuTe ecosystem 

Skills Required

  • MS, PhD, or equivalent experience in Computer Science, Software Engineering, or a related field
  • 2+ years of relevant work experience
  • Excellent programming skills in Python and strong proficiency in C++
  • Hands-on experience with DSLs, compilers, or code generation systems
  • Strong command of the MLIR/LLVM stack, including IR design and pass optimization
  • Strong communication skills and ability to thrive in a collaborative environment

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Ericsson Logo Ericsson

Hardware Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Beijing, CHN
88000 Employees

Ericsson Logo Ericsson

New Grad-Radio Developer TRX-BJ

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Beijing, CHN
88000 Employees

Ericsson Logo Ericsson

Network Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
2 Locations
88000 Employees

Ericsson Logo Ericsson

New Grad-Board Power Design-BJ

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Beijing, CHN
88000 Employees

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account