Senior ML Infrastructure Engineer

Posted 3 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Artificial Intelligence • Information Technology
The Role
The Senior ML Infrastructure Engineer will own and optimize GPU infrastructure for training tabular foundation models, ensuring high efficiency and collaboration with research teams.
Summary Generated by Built In
Who we are

Foundation models have transformed text and images, but structured data - the largest and most consequential data modality in the world - has remained untouched. Tables power every clinical trial, every financial model, every scientific experiment, every business decision. No one has built a foundation model that truly understands them.

Until now. What LLMs did for language, we're doing for tables.

Momentum: We pioneered tabular foundation models and are now the world-leading organization in structured data ML. Our TabPFN v2 model was published in Nature and set a new state-of-the-art for tabular machine learning. Since its release, we've scaled model capabilities more than 20x, reached 3M+ downloads, 6,000+ GitHub stars, and are seeing accelerating adoption across research and industry - from detecting lung disease with Oxford Cancer Analytics to preventing train failures with Hitachi to improving clinical trial decisions with BostonGene.

The hardest work is in front of us. We're scaling tabular foundation models to handle millions of rows, thousands of features, real-time inference, and entirely new data modalities - while building the infrastructure to deploy them in production across some of the most demanding industries on earth. These are open problems no one else is working on at this level.

Our team: We’re a small, highly selective team of 20+ engineers and researchers, selected from over 5,000 applicants, with backgrounds spanning Google, Apple, Amazon, Microsoft, G-Research, Jane Street, Goldman Sachs, and CERN, led by Frank Hutter, Noah Hollmann and Sauraj Gambhir and advised by world-leading AI researchers such as Bernhard Schölkopf and Turing Award winner Yann LeCun. We ship fast, create top-tier research, and hold each other to an extremely high bar.

What’s Next: In 2025, we raised €9m pre-seed led by Balderton Capital, backed by leaders from Hugging Face, DeepMind, and Black Forest Labs. The next modality shift in AI is happening - and we're hiring the team that makes it.

About the Role

We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures.

Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it.

What you'll work on:

  • Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization

  • Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs

  • Architect the next generation of our infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning against growing compute demands

  • Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling that keeps research iteration speed high

  • Own the compute budget. You understand cost per FLOP across providers and hardware, and you hate wasted compute

Tech stack: Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton

You may be a good fit if you have:

  • 5+ years building and operating production GPU infrastructure or distributed training systems at scale. At a major AI lab, a well-funded ML startup, or an HPC environment

  • Deep hands-on experience with Slurm and cluster management. You've debugged scheduling failures, optimized utilization across multi-tenant GPU workloads, and operated infrastructure where downtime has real cost

  • Expert-level systems thinking: memory bandwidth, GPU profiling. You reason about hardware, not configs

  • Strong Python and genuine fluency with PyTorch internals. Enough to profile a training run and tell whether the bottleneck is data loading, communication, or compute

  • Track record of making infrastructure decisions that measurably improved training throughput or cost efficiency

  • Strong AI tooling skills. You use Claude Code, Cursor, or similar fluently to move fast without sacrificing quality

Bonus:

  • Experience operating at tens-of-millions-scale GPU spend

  • Multi-cloud or hybrid HPC/cloud infrastructure experience

  • Triton, CUDA, or custom kernel experience

  • Experience scaling from single cluster to multi-cluster orchestration

  • Background building experiment tracking, model registry, or ML pipeline tooling

Our Commitments
  • We believe the best products and teams come from a wide range of perspectives, experiences, and backgrounds. That’s why we welcome applications from people of all identities and walks of life, especially anyone who’s ever felt discouraged by "not checking every box."

  • We’re committed to creating a safe, inclusive environment and providing equal opportunities regardless of gender, sexual orientation, origin, disabilities, or any other traits that make you who you are.

Top Skills

Docker
GCP
Github Actions
PyTorch
Slurm
Triton
Uv
Wandb
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Berlin
11 Employees
Year Founded: 2024

What We Do

Prior Labs is building breakthrough foundation models that understand spreadsheets and databases - the lifeblood of science and business. While foundation models have transformed text and images, tabular data has remained largely untouched. We're tackling this opportunity to revolutionize how we approach scientific discovery, medical research, financial modeling, and business intelligence. Backed by Balderton Capital, XTX Ventures, SAP Founder Hans Werner-Hector's Hector Foundation, Atlantic Labs, Galion.exe and top AI leaders such as Peter Sarlin, Guy Podjarny, Thomas Wolf, Ed Grefenstette, Robin Rombach, Christopher Lynch and Ash Kulkarni.

Similar Jobs

Rapid7 Logo Rapid7

Cybersecurity Advisor

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
Germany
2400 Employees

HERE Technologies Logo HERE Technologies

Senior Security Engineer

Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Hybrid
3 Locations
6000 Employees

Zeta Global Logo Zeta Global

Technical Product Manager

AdTech • Artificial Intelligence • Marketing Tech • Software • Analytics
Easy Apply
Hybrid
Berlin, DEU
2429 Employees

MongoDB Logo MongoDB

Senior Customer Success Manager

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
Berlin, DEU
5550 Employees

Similar Companies Hiring

Milestone Systems Thumbnail
Software • Security • Other • Big Data Analytics • Artificial Intelligence • Analytics
Lake Oswego, OR
1500 Employees
Idler Thumbnail
Artificial Intelligence
San Francisco, California
6 Employees
Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account