ML Research Engineer, Training

Reposted 6 Hours Ago
San Francisco, CA, USA
In-Office
Entry level
Artificial Intelligence • Hardware • Robotics • Automation
The Role
Build and optimize end-to-end machine learning training infrastructure for large-scale multimodal robot data. Responsibilities include distributed training, data ingestion and transformation, checkpointing, orchestration, experiment tracking, sampling and curation, performance profiling, debugging data and training issues, and converting research prototypes into production infrastructure. The role may also involve robot learning, cloud and cluster infrastructure, post-training, on-robot inference optimization, CUDA or Triton kernels, and video-heavy datasets.
Summary Generated by Built In
Join Us, and Ship Robots

Weave was founded to build the robots we’d want to have in our own home. We believe the next generation of robotics will transform everyday life by enabling people to do more and to reclaim time to spend on what’s important.

We also believe robots are in a sense like any other product: to matter, they have to ship. Our robots are already operating in real homes and businesses, giving us the opportunity to rapidly improve from real-world experience. With a growing team, strong customer demand, and capital for expansion, we’re entering an exciting stage of growth—and we’re looking for people with exceptional talent and standards to help bring home robotics to millions of households.

The Role

Most robot learning research is graded on evals that don't survive contact with the field. Ours is graded by robots doing useful work in real homes and businesses, every day. We're one of the first companies with a deployed fleet generating real-world robot data at terabyte scale. The pipeline and training stack you build are what turns that data into capability.

Model quality is set as much by training as by architecture: what data gets in, how it's sampled, whether the run is stable, or whether a silent bug ate the gradient three days ago. You'll own that layer from raw fleet uploads to the batch that hits the GPU. When the stack is right, ideas become models in training in days, and deployed in weeks.

Responsibilities
  • Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.

  • Large-scale data handling: Develop high-throughput data ingestion, transformation, and storage systems capable of processing terabytes of multimodal robot data, including video, proprioception, and sensor streams.

  • Optimize research productivity: Grow the codebase that makes experiments reproducible, scalable, and easy to launch, monitor, retry/recover, and debug.

  • Sampling and Curation: Drive sampling and curation decisions that show up in model behavior.

  • Make runs fast and honest: profile and fix throughput bottlenecks, chase down loss spikes and silent data bugs, keep results reproducible enough to trust: from data loading to GPU kernels.

  • From research to production: Turn research prototypes into infrastructure the whole team trains on.

What You’ll Bring
  • ML System Expertise: Deep PyTorch or JAX experience, including multi-node distributed training (FSDP, DDP, or equivalent) on real workloads.

  • Performance engineering: Experience profiling and optimizing GPU utilization, data pipelines, I/O bottlenecks, memory usage, and distributed training performance, including CUDA-level profiling tools (e.g. Nsight Systems) and NCCL tuning.

  • Training Run Judgement: You can read a loss curve, tell instability from a data bug and know when to kill a run.

Nice To Have
  • Robot learning exposure: you’ve trained policies (VLAs, world models, RL) and can tell a data problem from a model problem.

  • Cluster and cloud infrastructure experience: Kubernetes, SLURM, GCP/AWS.

  • Large-scale post-training experience: SFT, reward modeling, RL fine-tuning.

  • On-robot inference optimization experience: TensorRT, quantization, distillation.

  • CUDA or Triton kernel work.

  • Experience with video-heavy datasets: transcoding, chunking, and the storage/compute tradeoffs of training on video at scale.

Skills Required

  • Deep experience with PyTorch or JAX
  • Experience with multi-node distributed training using FSDP, DDP, or equivalent
  • Experience profiling and optimizing GPU utilization, data pipelines, I/O bottlenecks, memory usage, and distributed training performance
  • Experience with CUDA-level profiling tools such as Nsight Systems and NCCL tuning
  • Ability to diagnose training instability, loss spikes, silent data bugs, and determine when to stop a run
  • Experience training robot-learning policies, vision-language-action models, world models, or reinforcement-learning systems
  • Experience with Kubernetes, SLURM, GCP, or AWS
  • Large-scale post-training experience, including supervised fine-tuning, reward modeling, or reinforcement-learning fine-tuning
  • On-robot inference optimization experience with TensorRT, quantization, or distillation
  • CUDA or Triton kernel development experience
  • Experience with video-heavy datasets, including transcoding, chunking, and storage or compute optimization
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
16 Employees
Year Founded: 2024

What We Do

Weave Robotics is a San Francisco-based startup focused on developing practical, autonomous personal robots for the home. Their debut product, Isaac 0, is a stationary laundry-folding robot designed to save users time by autonomously tidying and folding clothes. The company aims to transition advanced robotics research into real-world home products, starting with laundry as a primary use case to return time to households.

Similar Jobs

Metamorphic Additive Manufacturing Ltd Logo Metamorphic Additive Manufacturing Ltd

ML Research Engineer (Distributed Training)

3D Printing • Consulting • Design • Manufacturing
In-Office
Palo Alto, CA, USA
200K-280K Annually

CoreWeave Logo CoreWeave

Systems Engineer

Cloud • Information Technology • Machine Learning
In-Office
5 Locations
1450 Employees
182K-242K Annually

Outset AI Logo Outset AI

Program Manager

Artificial Intelligence • Software
Hybrid
2 Locations
30 Employees
150K-200K Annually

ServiceNow Logo ServiceNow

Director, Insights and Measurement

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
221K-387K Annually

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account