Software Engineer, ML Serving - Rime Ai

Posted 2 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Entry level
Angel or VC Firm • Fintech
The Role
Own and architect Rime’s real-time TTS model-serving infrastructure, including GPU-backed inference, APIs, distributed serving, hardware compatibility, CI/CD, observability, on-call reliability, and GPU resource and cost management. The role requires cloud infrastructure, containerization, infrastructure-as-code, and distributed ML serving expertise, with preferred experience in training, streaming protocols, audio systems, multi-cloud environments, and configuration management.
Summary Generated by Built In

Rime is a foundation modeling company that builds voice AI for enterprises running customer experiences at scale. Our models are purpose-built for high-volume conversational deployments, engineered for the accuracy, performance, and deployment flexibility that production environments actually demand.

We started from a different premise than the rest of the field: build voice AI for human connection, not slop. Before we trained a single model, we built our own corpus: full-duplex, studio-quality conversational speech of normal people, recorded and annotated by linguists. It's why our models are unparalleled in naturalism, and it's why enterprises pick Rime when pilots need to make it to production.

 Role Overview

We're hiring a Software Engineer to own the serving infrastructure that connects Rime's inference engines to the world. This role sits at the intersection of ML systems and cloud infrastructure — you'll work directly on model inference and cloud infrastructure to build, harden, and scale the systems that stream voice at real-time latency. As Rime moves toward its next-generation architecture, you'll be a core architect of how our models get served.


What You'll Own

  • Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.

  • Model optimization from a single-node to disaggregated fleet serving.

  • Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.

  • Continuous integration and deployment workflows for the model serving pipeline.

  • Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.

  • Resource provision, cost management across our GPU fleet.

What We're Looking For

  • Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.

  • Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).

  • Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).

  • IaC experience — Terraform, Packer, or comparable. You should have opinions about how to do this right.

  • On-call is part of the job. You treat production reliability as a shared responsibility.

Nice to Have

  • Experience with multinode training (DDP, FSDP, etc.).

  • Experience with gRPC or other bidirectional binary streaming protocols.

  • Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).

  • Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.

  • Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).

  • Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).

  • SRE, DevOps, or platform engineering background at a startup.

  • Experience at an early-stage company.

Why Join Rime

  • Build the serving infrastructure behind a category-defining voice AI company from the ground up.

  • You will bring in experience that no one else currently has at the company: you can help us set the vision.

  • Direct collaboration with the inference, platform, and ML teams — no handoff culture.

  • The systems you build determine what experiences our customers can deploy at scale.

  • Meaningful equity upside at an early stage.

  • High ownership, high standards, low bureaucracy.

  • SF / Bay Area.

At Rime, we...

  • Are outliers

  • Cut through the hype to focus on the craft

  • Move fast with agency and freedom

  • Maintain a growth mindset, finding joy in the struggle

  • Do the right things, knowing that it'll lead to making money

  • If that sounds like you too, you'll be a great fit for Rime!

Skills Required

  • Hands-on experience with real-time multinode ML serving infrastructure
  • Experience with ML serving frameworks such as NVIDIA Dynamo, Triton, vLLM, SGLang, or equivalent
  • Experience with distributed or disaggregated model serving, including Tensor Parallel, Pipeline Parallel, or equivalent
  • Strong cloud infrastructure fundamentals, including Linux internals, networking, and containerization
  • Experience with Docker and Kubernetes
  • Infrastructure-as-code experience with Terraform, Packer, or comparable tools
  • Willingness to participate in an on-call rotation and share production reliability responsibility
  • Experience with multinode training, such as DDP or FSDP
  • Experience with gRPC or other bidirectional binary streaming protocols
  • Experience with audio streaming technologies such as WebRTC or WebSockets
  • Experience working in a multilingual monorepo
  • Experience with multi-cloud infrastructure such as AWS, GCP, or OCI
  • Familiarity with configuration management tools such as Ansible, Chef, or Puppet
  • SRE, DevOps, or platform engineering background at a startup
  • Experience at an early-stage company
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Menlo Park, CA
32 Employees
Year Founded: 2018

What We Do

Unusual Ventures is raising the bar for what founders should expect from their venture investors. We enable startups with the hands on support and expertise they need to be successful during their early stage journey. Unusual Ventures is an early-stage firm that invests in both enterprise and consumer startups.

Similar Jobs

Boeing Logo Boeing

Flight Software Engineers (Associate / Experienced / Senior)

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
El Segundo, CA, USA
170000 Employees
106K-232K Annually

Boeing Logo Boeing

Operations Center Fleet Monitoring Engineer (Associate or Experienced)

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Seal Beach, CA, USA
170000 Employees
99K-162K Annually

Boeing Logo Boeing

Integrated Product Team (IPT) Lead

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
El Segundo, CA, USA
170000 Employees
127K-212K Annually

Boeing Logo Boeing

Static Timing Analysis (STA) Engineer - (Lead or Senior)

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
El Segundo, CA, USA
170000 Employees
146K-239K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account