Software Engineer - ML Infrastructure

Posted 6 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Senior level
Healthtech • Software
The Role
Build and maintain ML infrastructure and data systems for medical imaging: distributed training, RL training stack, high-throughput data pipelines, centralized storage, and production serving to accelerate research-to-production iteration.
Summary Generated by Built In
About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

 
Role Overview

We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks.

 

Sitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production.

 
Key Responsibilities
  • Partner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change

  • Build a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.

  • Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.

  • Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.

  • Contribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team.

  • Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.

Qualifications
  • 6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production

  • Have 2+ years of experience building ML infrastructure or systems in production

  • Strong Python skills and expertise in PyTorch or JAX

  • Experience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research

  • Experience building infrastructure or platforms specifically for research or machine learning workflows

  • Deep experience building and operating Kubernetes and cloud infrastructure at scale

  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient

  • Prior experience as a technical lead or mentor for other engineers

Preferred Qualifications
  • Experience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy

  • Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems

  • Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production

  • Experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.

    • Experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.

  • Experience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows

  • Familiarity with vision-language models (VLMs) or multimodal architectures

The anticipated annual base salary for this position is up to $250,000. This range does not include any other compensation components or other benefits for which an individual may be eligible. The actual base salary offered depends on a variety of factors, which may include as applicable, the qualifications of the individual applicant for the position, years of relevant experience, specific and unique skills, level of education attained, certifications or other professional licenses held, and the location in which the applicant lives and/or from which they will be performing the job.

Skills Required

  • 5+ years building ML infrastructure, data pipelines, or ML systems in production
  • Strong Python skills
  • Expertise in PyTorch or JAX
  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism)
  • Hands-on experience with data pipeline technologies (e.g., Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk) and schema design
  • Experience with distributed systems and cloud infrastructure (AWS or GCP)
  • Experience with containerization (Docker and Kubernetes)
  • Track record of building scalable data systems and shipping production ML infrastructure
  • Ability to move quickly and handle competing priorities in a fast-paced environment
  • Experience building reinforcement learning training infrastructure (rollout generation, reward-model serving, experience collection)
  • Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton)
  • Experience building internal training or experimentation platforms used by research teams
  • Experience supporting A/B testing and experimentation workflows, canary deployments, and monitoring
  • Familiarity with vision-language models or multimodal architectures
  • Experience with medical imaging formats (DICOM) and healthcare data standards
  • Familiarity with MLOps practices and model deployment pipelines
  • Experience with privacy-preserving data systems and HIPAA compliance
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
6 Employees
Year Founded: 2024

What We Do

Epsilon Health exists to solve the looming global radiology crisis before it reshapes patient care. Radiology underpins nearly every medical specialty, yet it is one of the most strained parts of the healthcare system. We’re rethinking how imaging and interpretation work to make radiology faster, more reliable, and future-proof.

Similar Jobs

Snap Inc. Logo Snap Inc.

Principal Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
5 Locations
5000 Employees
235K-414K Annually

Claryo, Inc. Logo Claryo, Inc.

Senior Software Engineer

Artificial Intelligence • Logistics • Robotics • Software
In-Office
San Francisco, CA, USA
14 Employees
170K-190K Annually

Snap Inc. Logo Snap Inc.

Staff Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
3 Locations
5000 Employees
195K-343K Annually

Handshake Logo Handshake

Senior Software Engineer

Edtech • Enterprise Web • HR Tech • Software
In-Office
San Francisco, CA, USA
700 Employees
176K-220K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account