Software Engineer - ML Infrastructure

Posted 8 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
250K-300K Annually
Senior level
Professional Services • Consulting
The Role
Own machine learning infrastructure spanning distributed training, reinforcement learning, data pipelines, model serving, deployment, monitoring, and rollout tooling. Partner with researchers to productionize experimental workflows for clinical AI applications. The role also involves GPU scheduling, autoscaling, reproducible environments, online-learning loops, checkpointing, evaluation, and technical leadership or mentoring.
Summary Generated by Built In

About the company

Our client is a healthcare technology company developing AI-enabled imaging tools. Researchers and engineers work together on production machine learning systems for clinical applications.

The role

Raydar is recruiting for this opportunity through the Paraform network. The position is with our client. Own ML infrastructure across distributed training, reinforcement learning, and production serving. You will establish engineering practices and help researchers translate experiments into reliable systems.

What you'll do

- Build distributed training infrastructure with parallelism and checkpointing.

- Develop reinforcement-learning infrastructure for rollout generation, reward models, and experience collection.

- Partner with researchers to turn experimental workflows into production systems.

- Build data loading and preprocessing pipelines for multimodal datasets.

- Improve model serving, canary deployment, monitoring, and rollout tooling.


Requirements

What we're looking for

- 6+ years of ML infrastructure or distributed systems experience.

- Strong Python and hands-on production ML serving experience.

- Kubernetes and Docker experience with GPU scheduling, autoscaling, and reproducible environments.

- Distributed training expertise in PyTorch or JAX, including FSDP, DeepSpeed, or comparable approaches.

- Experience with RL or online-learning loops, logging, checkpointing, and evaluation.

- Breadth across the ML infrastructure stack plus technical leadership or mentoring experience.

Bonus points

- Startup experience.

- A/B testing and production ML experimentation platforms.

- A computer science or other STEM degree.


Benefits

Compensation and benefits

- Base salary: USD 250,000 to 300,000 per year.

- Equity: Competitive equity.

Location and work model

- San Francisco, California, United States.

- Five days per week onsite.

- Visa transfers and new visa sponsorships are supported.

Skills Required

  • 6+ years of ML infrastructure or distributed systems experience
  • Strong Python experience
  • Hands-on production machine learning serving experience
  • Kubernetes and Docker experience, including GPU scheduling, autoscaling, and reproducible environments
  • Distributed training expertise in PyTorch, JAX, FSDP, DeepSpeed, or comparable approaches
  • Experience with reinforcement learning or online-learning loops, logging, checkpointing, and evaluation
  • Breadth across the machine learning infrastructure stack
  • Technical leadership or mentoring experience
  • Startup experience
  • Experience with A/B testing and production machine learning experimentation platforms
  • Computer science or other STEM degree
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
28 Employees
Year Founded: 2021

What We Do

Raydar is a talent acquisition and business consulting firm that connects world-class and emerging talent with growing organizations. It supports companies through team development, strategic hiring, and customized growth solutions, helping clients recruit roles such as engineers, product managers, executives, legal counsel, and quantitative traders. Raydar focuses on understanding each organization’s needs, culture, and long-term goals to build high-impact teams.

Similar Jobs

Snap Inc. Logo Snap Inc.

Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
3 Locations
5000 Employees
178K-313K Annually
Hybrid
2 Locations
289097 Employees

Snap Inc. Logo Snap Inc.

Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
Palo Alto, CA, USA
5000 Employees
133K-235K Annually

Watney Robotics Inc Logo Watney Robotics Inc

Staff Software Engineer

Artificial Intelligence • Information Technology • Robotics • Software
In-Office
San Francisco, CA, USA
13 Employees

Similar Companies Hiring

Fora Thumbnail
Agency • On-Demand • Professional Services • Sales • Software • Travel • Hospitality
New York, NY
250 Employees
Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
150 Employees
Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account