Machine Learning Engineer, Ops

Posted Yesterday
Be an Early Applicant
27 Locations
Remote
125K-165K Annually
Mid level
Artificial Intelligence • Software
The Role
Design, deploy, and scale low-latency inference infrastructure for generative audio models (TTS, ASR, voice conversion). Build high-performance inference engines, Kubernetes-based autoscaling, CI/CD pipelines, observability, and GPU optimization to bridge research and production for streaming and batch workloads.
Summary Generated by Built In

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

 

About the Role:

We are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.

What You’ll Do:

  • Design and maintain inference infrastructure for generative audio model architectures.

  • Implement and manage high-performance inference engines.

  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.

  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.

  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.

  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.

  • Optimize inference performance for both streaming and batch applications.

What You’ll Bring:

  • Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.

  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.

  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.

  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.

  • Experience with GPU-accelerated inference and performance profiling techniques.

  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Compensation:

The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

 

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Skills Required

  • Deep understanding of modern audio model architectures (TTS, ASR) and their inference requirements.
  • Hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production.
  • Background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
  • Experience with GPU-accelerated inference and performance profiling techniques.
  • Experience implementing and managing high-performance inference engines.
  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni).
  • Experience optimizing inference performance for both streaming and batch applications.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
364 Employees
Year Founded: 2023

What We Do

Cantina Labs, founded by Sean Parker, is a new social platform with the most advanced AI character creator. Build, share, and interact with AI bots and your friends directly in the Cantina or across the internet. Cantina bots are lifelike, social creatures, capable of interacting wherever humans go on the internet. Recreate yourself using powerful AI, imagine someone new, or choose from thousands of existing characters. Bots are a new media type that offer a way for creators to share infinitely scalable and personalized content experiences combined with seamless group chat across voice, video, and text.

Similar Jobs

Pragmatike Logo Pragmatike

Principal ML Ops Engineer (EMEA Remote)

Information Technology • Software
In-Office or Remote
20 Locations
11 Employees

Pragmatike Logo Pragmatike

ML Ops Engineer (EMEA Remote)

Information Technology • Software
In-Office or Remote
20 Locations
11 Employees

Tufin Logo Tufin

Consultant

Security • Cybersecurity
Remote or Hybrid
27 Locations
500 Employees

Pfizer Logo Pfizer

Audit Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote
29 Locations
121990 Employees
163K-272K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account