Compute Infrastructure Lead

Posted 12 Days Ago
Be an Early Applicant
Paris, Île-de-France, FRA
In-Office
Senior level
Blockchain • Fintech • Software • Financial Services
The Role
Own and scale UMA’s compute infrastructure for large-scale AI training, evaluation, and data-processing workloads. Responsibilities include multi-provider GPU capacity, cloud scheduling, distributed training frameworks, checkpointing, preemption, virtualized development environments, observability, reliability, cost optimization, and provider relationships. The role requires hands-on operation of production-grade GPU clusters and may grow into leadership of the compute infrastructure team.
Summary Generated by Built In
Your Mission

As Compute Infrastructure Lead, you will own and scale the compute backbone of UMA: the systems that provision, schedule, and run training, evaluation, and data-processing workloads — reliably, efficiently, and at scale — so our models can go from research to production without the cluster becoming the bottleneck.

This is a hands-on, high-impact role. We already train on a dedicated GPU cluster with a working training stack and a strong team behind it, so you won't be starting from zero — but you'll have the mandate to shape the architecture that takes us from a research cluster to a production-scale, multi-provider fleet, and to production-grade reliability as we start deploying POCs with industry partners. You'll take ownership of the compute platform end to end — multi-provider capacity, scheduling, distributed training / eval / processing, virtualized developer environments, observability, cost, and the tooling researchers and engineers actually use — and, if that's where you want to go, grow into leading the compute infrastructure team as it scales.

The technical problem is unusually rich for this stage. We are de-risking a stack built on pre-training and online RL, then industrializing it: heterogeneous hardware (training GPUs, cheaper eval and processing GPUs, CPU), a real-time learning loop, orchestrated data processing, interactive VMs on the cluster, and a fleet that will grow fast. Much of what we need — a true multi-provider compute fabric with elastic scheduling, dynamic checkpointing, unified observability, and jobs that resume themselves — does not exist off the shelf. Data infrastructure (datasets, storage, versioning) is owned by a sister role; this role is compute, including how processing jobs actually run on it.

Key responsibilities :

  • Own our compute platform end to end — from provisioning GPU capacity across cloud providers to keeping training, eval, and processing jobs running at high utilization, with reliability, cost, and researcher velocity as first-class goals

  • Build a multi-provider management layer so we can place, burst, and fail over workloads across GPU clouds, hyperscalers, and HPC without rewriting jobs

  • Design and operate the cloud scheduler — quotas, priority, preemption, topology-aware placement, and dynamic checkpointing so jobs survive node failure, preemption, and provider switches

  • Stand up a distributed compute framework for training and evaluation on heterogeneous hardware (e.g. Ray / similar), including the real-time / online-learning path

  • Orchestrate data-processing workloads at scale — CPU and cheaper GPUs, batch and streaming — so post-processing, dataset jobs, and training share one reliable compute fabric instead of ad-hoc scripts

  • Deliver virtualized GPU/CPU dev sessions (VMs on the cluster) so engineers iterate interactively on the same hardware and software they train on, without burning dedicated boxes

  • Build observability that works from any provider — system metrics (Prometheus, Grafana), job traces and logs, and model metrics (e.g. MLflow) — so a hung NCCL job, a silent GPU, or a broken training curve is diagnosable in minutes, not days

  • Own capacity, cost, and provider relationships as a technical lead: forecast demand, pick the right mix of hardware and contracts, and help negotiate pricing and terms. This is not a commercial role — but procurement is part of making the infra succeed

  • Help set production-grade practices (testing, reliability, fast iteration) as we move from R&D to partner POCs, and grow into leading the compute infra team if that's the path you want

What You Bring to the Table
  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering, at a senior, lead, or staff level

  • Proven track record building and operating infrastructure for large-scale AI model training — not inference-only. Multi-node GPU clusters, distributed training (PyTorch / NCCL or equivalent), and keeping long-running jobs healthy at scale

  • Deep, hands-on experience with GPU clouds and cluster operations: provisioning, Linux, high-performance networking (InfiniBand / RoCE), storage for training, utilization, and GPU/node failure modes

  • Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot, or similar) — including checkpointing, elasticity, and preemption — so GPUs stay busy and jobs come back from failure

  • Treat observability, reliability, and cost as core engineering concerns, not afterthoughts

  • Strong Python and systems engineering, with the taste to build tooling that researchers actually want to use

  • Experience working with GPU providers on capacity and commercial terms — you can read a contract, push on price and availability, and still be the person who debugs the cluster at 2am

  • Ability to reason about systems end-to-end — performance, scalability, reliability, cost — and make and defend the right trade-offs

  • Thrive in a hands-on, fast-paced startup, building from a real (but small) cluster toward a production fleet: autonomous, rigorous, execution-driven, easy to work with, and broadly curious about AI and systems

  • Bonus : online / continuous RL, real-time training loops, or other always-on learning systems

  • Bonus : multi-cloud / multi-provider fabrics (SkyPilot, similar), Ray / Anyscale, HPC centers, VM-based GPU workstations / interactive cluster sessions, or standing up clusters from tens to hundreds of nodes

  • Bonus : robotics, autonomous vehicles, or other embodied/physical-AI training stacks — adjacent large-scale training (LLMs, multimodal, AV) counts strongly; robotics itself is not required

  • Bonus : public projects, open-source contributions, maintained tools, or technical writing

  • We value exceptional builders over perfect resumes. If you have a world-class track record building training infrastructure at scale and the drive to build the compute backbone that lets a robotics company scale, we strongly encourage you to apply — even if you don't tick every box. Robotics experience is a plus, not a requirement.

Skills Required

  • 8+ years of experience in ML/compute infrastructure, HPC, or large-scale GPU platform engineering at a senior, lead, or staff level
  • Experience building and operating infrastructure for large-scale AI model training, including multi-node GPU clusters and distributed training
  • Experience with PyTorch, NCCL, or equivalent distributed training technologies
  • Hands-on experience with GPU clouds and cluster operations, including provisioning, Linux, high-performance networking, training storage, utilization, and GPU or node failure modes
  • Experience building or owning schedulers and distributed frameworks such as Slurm, Kubernetes, Ray, or SkyPilot
  • Experience with checkpointing, elasticity, and preemption
  • Strong Python and systems engineering skills
  • Experience working with GPU providers on capacity and commercial terms
  • Ability to reason about performance, scalability, reliability, and cost trade-offs end to end
  • Ability to work hands-on in a fast-paced startup environment
  • Experience with online or continuous reinforcement learning and real-time training loops
  • Experience with multi-cloud or multi-provider compute fabrics, Ray or Anyscale, HPC centers, VM-based GPU workstations, or interactive cluster sessions
  • Experience standing up clusters ranging from tens to hundreds of nodes
  • Robotics, autonomous vehicle, embodied AI, LLM, multimodal, or adjacent large-scale training experience
  • Public projects, open-source contributions, maintained tools, or technical writing
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
27 Employees
Year Founded: 2025

What We Do

We build general-purpose mobile and humanoid robots capable of human-level dexterity and understanding of the physical world. It will enable people to focus on what truly matters in their lives. We are based in Paris, FR. Join us: https://app.dover.com/jobs/uma

Similar Jobs

Mirakl Logo Mirakl

Accountant

eCommerce • Information Technology • Retail • Software
Hybrid
Paris, Île-de-France, FRA
750 Employees

Pfizer Logo Pfizer

Staff Software Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Serris, Seine-et-Marne, Île-de-France, FRA
16000 Employees
23K-31K Annually

MongoDB Logo MongoDB

Solutions Architect

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
Paris, Île-de-France, FRA
5550 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account