Helix AI Engineer, Training Infrastructure

Reposted 19 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
150K-350K Annually
Mid level
Artificial Intelligence • Robotics • Automation • Manufacturing
The Role
The role involves managing training infrastructure, implementing distributed training algorithms, and collaborating with AI researchers on model training.
Summary Generated by Built In
Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.
Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced Training Infrastructure Engineer to take our infrastructure to the next level. This role is focused on managing the training cluster, implementing distributed training algorithms, data loaders, and developer tools for AI researchers.
Responsibilities
  • Design, deploy, and maintain Figure's training clusters
  • Architect, optimize, and maintain scalable deep learning frameworks for training on massive robot datasets
  • Work together with AI researchers to implement training of new model architectures at a large scale
  • Implement distributed training, advanced parallelization strategies, and high-performance data loaders to reduce model development cycles
  • Profile, identify, and eliminate training bottlenecks at the hardware and software levels to maximize Model FLOPs Utilization (MFU)
  • Implement tooling for data processing, model experimentation, and continuous integration
Requirements
  • Strong software engineering fundamentals
  • Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related field
  • Extensive professional experience with Python and PyTorch
  • Proven track record of scaling and running large-scale training experiments personally on 800+ GPUs
  • Experience managing HPC clusters for deep neural network training
  • Minimum of 4 years of professional, full-time experience building reliable backend systems and infrastructure
Bonus Qualifications
  • Experience contributing to or maintaining open-source distributed training frameworks (Megatron-LM, DeepSpeed, TorchTitan)
  • Experience managing cloud infrastructure (AWS, Azure, GCP)
  • Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)
  • Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.)
  • Deep understanding of CUDA and hands-on experience writing custom GPU kernels to optimize training

The US base salary range for this full-time position is between $200,000 - $400,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Skills Required

  • Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or related field
  • Experience with Python and PyTorch
  • Experience managing HPC clusters for deep neural network training
  • Minimum of 4 years of professional, full-time experience building reliable backend systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
86 Employees
Year Founded: 2022

What We Do

Figure is an AI Robotics company building the world's first commercially viable autonomous humanoid robot. We are based in Sunnyvale, CA.

Similar Jobs

Adyen Logo Adyen

Senior Product Manager

Fintech • Payments • Financial Services
Easy Apply
Hybrid
San Francisco, CA, USA
4771 Employees
202K-302K Annually

LogicMonitor Logo LogicMonitor

Sr. Forward Deployed Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
San Francisco, CA, USA
1100 Employees
131K-175K Annually

Coupa Logo Coupa

Principal Software Engineer

Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
In-Office or Remote
San Francisco, CA, USA
3000 Employees
201K-281K Annually

Altana Logo Altana

Principal Product Manager

Artificial Intelligence • Machine Learning • Software
Easy Apply
Hybrid
4 Locations
228 Employees
215K-270K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
Chicago, Illinois
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account