Principal ML Performance Engineer (GPU Optimization)

Posted 4 Days Ago
3 Locations
In-Office or Remote
Senior level
Information Technology • Consulting
The Role
Optimize training and inference performance for structural and generative ML models. Develop custom CUDA and Triton kernels, use ML compilers, scale distributed training across multi-node GPU clusters, reduce inference costs, improve GPU utilization on GCP, and build benchmarking and profiling tools. Provide technical direction, select infrastructure, influence research teams, and mentor engineers.
Summary Generated by Built In
Principal ML Performance Engineer (GPU Optimization)

About Proxima

Proxima is a frontier AI and data generation company discovering the next generation of proximity therapeutics by making protein interactions programmable. Our platform brings together foundation-model machine learning, a scalable data generation engine, and a partnership track record exceeding $5B in collaborations across the world’s leading biopharma and tech organizations. We’ve recently closed an oversubscribed seed round with an elite group of VCs including DCVC, NVIDIA’s NVentures, AIX, Yosemite among others.

Neo-1 is our all-atom foundation model that combines state-of-the-art structure prediction and molecular generation in a single system. Neo-1 enables rapid exploration of chemical and structural space for high value, previously intractable targets, and in particular unlocks small molecule proximity therapeutics like molecular glues with AI for the first time.

In parallel, we are developing an advanced structural interactomics platform built on proprietary XLMS technology and a lab equipped with next-generation mass spectrometry instrumentation. This platform produces proteome-scale maps of protein interactions and helps identify small molecules that modulate proximity. Together with Neo-1, it creates an integrated system capable of co-folding protein complexes while generating candidate small molecules to influence those interactions.
Proximity-based therapeutics represent one of the most promising frontiers in modern drug discovery with the potential to treat previously intractable diseases and target ‘undruggable’ proteins. We’re building the tech and the team to make that happen. Come join us!

What you'll do
  • Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning

  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA) when beneficial

  • Scale distributed training across 32-64 nodes, employing FSDP, DeepSpeed, tensor and pipeline parallelism, and mixed precision

  • Reduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs

  • Manage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting

  • Develop benchmarks and profiling tools for the research team

What we need
  • Minimum of 6+ years experience in ML systems, HPC, or performance engineering, with a BS/MS/PhD in CS, EE, or related field

  • Demonstrated ability to set technical direction beyond coding: selecting infrastructure, influencing research teams, and mentoring engineers

  • Deep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks

  • Experience with CUDA and Triton, skilled at reading Nsight output, and strong understanding of memory bandwidth and occupancy

  • Experience with distributed training at multi-node scale

  • Strong proficiency in Python and C++

  • Able to name a model they made materially faster and quantify the improvement

Nice to haves
  • Experience in geometric deep learning, equivariant networks, or protein structure models such as AlphaFold, ESM, or RFdiffusion

  • Experience writing kernels for structure-model primitives, including triangle attention, triangle multiplicative updates, cuEquivariance, or FlashAttention for pair bias

  • Experience orchestrating large batch inference and managing Kubernetes GPU scheduling


Skills Required

  • 6+ years of experience in ML systems, HPC, or performance engineering
  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field
  • Ability to set technical direction, select infrastructure, influence research teams, and mentor engineers
  • Deep knowledge of PyTorch internals and experience profiling and resolving performance bottlenecks
  • Experience with CUDA and Triton
  • Ability to read Nsight output and understand memory bandwidth and occupancy
  • Experience with distributed training at multi-node scale
  • Strong proficiency in Python and C++
  • Ability to quantify material performance improvements to a model
  • Experience in geometric deep learning, equivariant networks, or protein structure models such as AlphaFold, ESM, or RFdiffusion
  • Experience writing kernels for structure-model primitives, cuEquivariance, or FlashAttention for pair bias
  • Experience orchestrating large-batch inference and managing Kubernetes GPU scheduling
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
London, England
500 Employees
Year Founded: 1994

What We Do

PROXIMA IS A WORLD-LEADING PROCUREMENT AND SUPPLY CHAIN CONSULTANCY We work alongside some of the world’s largest and most successful businesses to help them spend their money wisely and deliver purposeful and profitable change. We do this through an extensive suite of procurement consultancy services focused on cost optimization, organizational transformation, supply chain sustainability, and decarbonization. We are famous for our delivery as experienced specialists immersed in businesses, accelerating outcomes. We are proud to be part of Bain & Company - the leading management consultancy. Together, we deliver a set of integrated end-to-end procurement, supply chain, and supply chain sustainability offerings.

Similar Jobs

Pfizer Logo Pfizer

Senior Manager, HTA, Value and Evidence (HV&E), Genitourinary Cancer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
139K-232K Annually

Pfizer Logo Pfizer

Staff Software Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees

Cloudflare Logo Cloudflare

Senior Systems Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
8 Locations
4400 Employees
66K-91K Annually

Pfizer Logo Pfizer

Sustainability Senior Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
30 Locations
121990 Employees
112K-207K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account