ML Engineer - Infrastructure

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Artificial Intelligence • Robotics • Software
The Role
Own the design, provisioning, operation, and optimization of cloud-based GPU clusters supporting large-scale model training and inference. Partner with AI engineers to improve training performance and hardware utilization, contribute to ML libraries and infrastructure tooling, develop GPU compute strategies, optimize multi-cloud capacity and costs, and improve reliability, testing, documentation, and engineering practices.
Summary Generated by Built In
About Flexion

At Flexion, we are building the autonomy stack for humanoid robots. Our mission is to drive the transition from fragile prototypes to real-world deployments of humanoids. We were founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich) and backed by leading international VC firms. In just months, we went from our first line of code to deploying real humanoid capabilities with our customers, leveraging simulation and reinforcement learning. Today, we are rapidly expanding the capabilities of our autonomy stack, our customer base, and our team.

The role

We are looking for an experienced ML engineer to join Flexion’s experienced infrastructure team and take ownership of Flexion’s GPU compute platforms. This is a senior, on-site role with significant scope. 

At Flexion, we are building the brain for humanoid robots, which involves training foundation models with vast amounts of data on large GPU clusters. You will own the design, bring-up, operation and optimization of performant clusters. You will work with AI engineers to help them optimize their training speed and hardware utilization. You will also influence strategic compute planning and contribute to new tools and platforms for iterating on our AI models efficiently. This will put you at the heart of Flexion’s AI development and allow you to directly impact the execution of our ambitious roadmap. You will closely collaborate with the company’s leadership, engineers of the infrastructure team and AI engineers across the company.

Key responsibilities
  • Architect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently. Work on cluster provisioning, job schedulers and monitoring systems.
  • Help AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries.
  • Contribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them. 
  • Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs.
  • Raise the bar on engineering practices, including testing, code quality, documentation, and system reliability.

Requirements
  • Degree in Computer Science, Electrical Engineering or Software Engineering (or equivalent practical experience) plus significant industry experience.
  • Hands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware. This can include bringing up and running the cluster, writing and optimizing training/inference code, building ML pipelines, etc.
  • Proficiency in Python and working knowledge of PyTorch.
  • Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
  • Experience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructure.
  • Experience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.

Nice-to-haves

  • Familiarity with profilers (e.g., PyTorch Profiler, Dynolog, HTA, Nsight).
  • Experience with high-performance or parallel file systems (e.g., Lustre).
  • Experience provisioning compute on multiple cloud providers.
  • Experience with infrastructure-as-code and configuration management (Terraform, Ansible).

Benefits
  • Competitive Compensation
  • Joining a leading robotics team & exposure to never-done-before research
  • Energetic, collaborative culture with a bias for action and regular community events

Zurich

  • Enhanced pension plan
  • Relocation & permit sponsorship
  • Enhanced holiday & paid leave perks
  • Central Zürich office with top-tier robotics testing facilities and infrastructure

San Franciso

  • 401(k) with company contributions
  • Health, dental & vision coverage with the flexibility to choose your own plan
  • Open PTO policy & paid company holidays

Skills Required

  • Degree in Computer Science, Electrical Engineering, Software Engineering, or equivalent practical experience, plus significant industry experience
  • Hands-on experience training or deploying inference for billion-parameter models on distributed, multi-node GPU hardware
  • Proficiency in Python and working knowledge of PyTorch
  • Deep understanding of distributed training concepts including DDP, FSDP, and NCCL
  • Experience with at least one cloud platform such as AWS, GCP, Azure, or neoclouds, or with large-scale on-premises GPU infrastructure
  • Experience with Slurm and/or Kubernetes/KubeRay for job scheduling and orchestration
  • Familiarity with PyTorch Profiler, Dynolog, HTA, or Nsight
  • Experience with high-performance or parallel file systems such as Lustre
  • Experience provisioning compute across multiple cloud providers
  • Experience with infrastructure as code and configuration management using Terraform or Ansible
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Zürich
18 Employees

What We Do

We at Flexion Robotics (flexion.ai) are a young company in Zurich working on the next generation of humanoid robot software to enable robots to perform useful tasks autonomously. We work dynamically and move fast. The team is still fairly small and every new employee at this stage will have significant ownership of their current project.

Similar Jobs

Nebius Logo Nebius

Infrastructure Engineer

Artificial Intelligence • Information Technology • Consulting
In-Office or Remote
27 Locations
473 Employees

Skydio Logo Skydio

Autonomy Engineer - ML & DL Infrastructure

Artificial Intelligence • Hardware • Robotics • Software
Hybrid
Zürich, CHE
250 Employees

Benchling Logo Benchling

Consultant

Cloud • Healthtech • Social Impact • Software • Biotech
Hybrid
Zürich, CHE
605 Employees

Circle (circle.so) Logo Circle (circle.so)

Lead Engineer, AI Platform

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
43 Locations
250 Employees

Similar Companies Hiring

Unusual Machines, Inc. Thumbnail
Hardware • Robotics
Orlando, FL
190 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account