MLOps Engineer

Posted 5 Hours Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
145K-180K Annually
Entry level
Artificial Intelligence • eCommerce • Fashion • Retail • Software
The Role
Own the end-to-end ML lifecycle platform, including training-as-a-service, experiment tracking, model CI/CD, data versioning, monitoring, and cost governance. Build infrastructure for multi-GPU training, automated evaluation gates, canary releases, A/B testing, and checkpoint management. Partner with ML Scientists to improve reproducibility and workflow efficiency, integrate external model providers, and help guide platform direction while mentoring engineers.
Summary Generated by Built In

SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences. Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.

About the team

AI Platform scales SPREEAI's infrastructure: productionizing ML model checkpoints, running the API services behind the Partner Portal, optimizing inference serving, and giving ML Scientists training-as-a-service so they can iterate without managing infrastructure themselves.

About the role

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

What you'll do
  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale
  • Partner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions
  • Evaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform
What you'll bring
  • Experience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)
  • Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment
You'll thrive here if

You treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident. You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.

Skills Required

  • Experience building or operating an ML platform
  • Experience with training orchestration using Ray, Kubeflow, Airflow, or a custom solution
  • Experience with Docker
  • Experience with Kubernetes, including Jobs, CronJobs, and Helm
  • Proficiency in Python or Go for pipeline orchestration and infrastructure tooling
  • Production-scale experience with experiment tracking and data versioning tools
  • Ability to own platform direction and mentor engineers
  • Ability to work across the ML lifecycle in an early-stage environment
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Incline Village, NV
35 Employees

What We Do

SPREEAI is redefining how the world shops—powered by AI, built for the future.

Similar Jobs

Gallatin AI, Inc. Logo Gallatin AI, Inc.

Machine Learning Operations (MLOps) Engineer

Artificial Intelligence • Logistics • Software • Defense
In-Office
3 Locations
45 Employees
80K-210K Annually
In-Office or Remote
2 Locations
300 Employees
152K-230K Annually
Hybrid
San Diego, CA, USA
955 Employees
106K-189K Annually
Hybrid
2 Locations
289097 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account