Staff Backend Engineer, Ambient AI

Posted 7 Hours Ago
Be an Early Applicant
2 Locations
Hybrid
210K-275K Annually
Senior level
Information Technology • Software
The Role
Design, build, and operate backend infrastructure for ambient healthcare AI, including audio ingestion, inference orchestration, distributed workflows, APIs, storage, observability, and cloud systems. Improve reliability, scalability, latency, and cost across real-time and asynchronous workloads. Lead incident response, establish SLOs, develop recovery and debugging tools, partner across engineering and clinical teams, and mentor engineers on distributed systems and operational practices.
Summary Generated by Built In

At Commure, we're building the AI Operating System for healthcare, the foundation that defines how care is delivered, documented, and financed. Our platform spans the full care journey: Ambient AI and Dictation eliminating documentation burden at the point of care, intelligent Agents automating patient and revenue workflows, and autonomous RCM processing billions in claims, all on a single AI-native platform integrated with 60+ EHRs.

Healthcare carries a $1 trillion administrative burden and we're at the center of transforming it. Today, 500,000+ clinicians across 500+ healthcare organizations nationwide trust Commure to handle $25B+ in annual claims and support over 200 million patient interactions. Our latest $70M raise at a $7B valuation reflects the confidence the market has placed in this mission. We've also been named to the Fortune Future 50 list and the 2026 AI Breakthrough Awards for “Overall NLP Company of the Year.”

Our team works directly alongside clinicians, not through layers of process, which means the gap between what you build and its impact on patient care is immediate. We move fast, deploy daily, and take full ownership from early thinking to production. If you're energized by hard problems, high stakes, and a team that holds itself to a high bar, you'll find your people here.

The future of healthcare is being built right now. Come deliver this transformation.

About the Role

We’re building a next-generation ambient AI product suite for healthcare, helping clinicians focus on patient care by automatically capturing, understanding, and structuring clinical conversations.

Behind that experience is a distributed system responsible for ingesting long-form audio, processing real-time and asynchronous streams, orchestrating multiple AI models, and reliably producing clinical outputs. These workflows must remain observable and recoverable across network interruptions, partial failures, model timeouts, and rapidly changing inference infrastructure.

We’re hiring a Staff Backend Engineer to help build and operate this foundation. You’ll work across audio ingestion, media processing, inference orchestration, distributed workflows, storage, observability, and cloud infrastructure.

This role is ideal for an engineer who enjoys making complex systems dependable. You think carefully about failure modes, idempotency, backpressure, data integrity, latency, and operational simplicity. You also understand that infrastructure exists to serve the product: the systems you build must ultimately create a fast, reliable, and trustworthy experience for clinicians.

What You’ll Do

  • Design, build, and operate backend systems that power our Ambient AI products across mobile and web.

  • Build and evolve inference workflows that coordinate transcription, diarization, language models, clinical extraction, summarization, and other AI capabilities.

  • Develop orchestration systems for long-running, multi-stage workflows, including:

    • Scheduling, queueing, and workload prioritization

    • Retries, timeouts, fallbacks, and dead-letter handling

    • Idempotency, replay, and safe workflow recovery

    • Model routing, versioning, and configuration

    • Progress tracking and user-visible workflow state

    • Graceful degradation when dependencies are unavailable

  • Improve the reliability and scalability of services operating under variable, compute-intensive workloads.

  • Define and maintain service-level objectives for critical workflows, including availability, processing latency, completion rates, and data durability.

  • Build strong observability across services and pipelines through structured logging, metrics, distributed tracing, dashboards, alerting, and diagnostic tooling.

  • Lead incident response and post-incident analysis for important production failures, turning operational lessons into lasting improvements.

  • Identify systemic failure patterns and eliminate them through better abstractions, automation, testing, and architecture.

  • Improve cloud infrastructure, deployment systems, capacity planning, and operational tooling as the platform scales.

  • Design APIs and data models that allow mobile, web, and internal systems to interact with long-running workflows safely and predictably.

  • Build tools that make inference pipelines easier to inspect, test, replay, evaluate, and debug.

  • Balance reliability, latency, quality, and infrastructure cost across AI and media-processing workloads.

  • Partner closely with mobile, product, AI, security, and clinical teams to turn new capabilities into production-ready workflows.

  • Mentor other engineers and help establish strong practices for distributed systems, observability, operational readiness, and backend architecture.

What You Have

  • 8+ years of professional backend or infrastructure engineering experience.

  • Experience designing, building, and operating production distributed systems.

  • Strong proficiency in at least one modern backend programming language.

  • Experience with asynchronous processing, message queues, event-driven systems, or durable workflow execution.

  • A strong understanding of distributed-systems concepts such as idempotency, consistency, concurrency, backpressure, retries, failure isolation, and eventual completion.

  • Experience operating services in a cloud environment, including deployment, monitoring, scaling, and incident response.

  • Experience designing APIs, service boundaries, and data models for complex product workflows.

  • A track record of improving system reliability, observability, scalability, or operational efficiency.

  • Strong debugging skills across application, infrastructure, data, and external dependency boundaries.

  • An ability to reason about both real-time and long-running workloads with different latency and durability requirements.

  • Comfort working across backend, infrastructure, product, and AI systems rather than within a narrowly defined layer.

  • An ownership mindset and bias for action. You identify important risks, create clarity, and drive ambiguous infrastructure work through completion.

Nice to Haves

  • Experience building speech-to-text, diarization, transcription, or other audio-based machine-learning workflows.

  • Experience orchestrating large language models or multi-model inference pipelines in production.

  • Familiarity with workflow orchestration systems such as Temporal or similar durable execution platforms.

  • Experience with event-streaming and messaging technologies such as Kafka, Pub/Sub, SQS, or similar systems.

  • Experience with model routing, inference gateways, rate limiting, batching, caching, or GPU-backed workloads.

  • Experience building internal tooling for workflow inspection, replay, evaluation, or model debugging.

  • Experience defining SLOs, managing error budgets, designing alerts, and leading production incident response.

  • Familiarity with distributed tracing and observability platforms such as Grafana, OpenTelemetry, Sentry, or similar tools.

  • Experience with containers, Kubernetes, infrastructure as code, and cloud-native deployment systems.

  • Experience optimizing systems for throughput, tail latency, infrastructure cost, and workload isolation.

  • Experience with offline clients, background synchronization, or systems that reconcile delayed and duplicated events.

  • Exposure to healthcare, HIPAA, SOC 2, encryption, privacy, security, or other regulated environments.

  • Experience supporting AI-native, real-time, or agentic product experiences.

Why Join

  • Build the backend foundation for an AI healthcare product used in real clinical workflows.

  • Solve challenging systems problems involving distributed orchestration, and production AI.

  • Make reliability improvements that clinicians experience directly through faster, safer, and more dependable workflows.

  • Shape how inference systems are orchestrated, observed, recovered, and scaled in production.

  • Work closely with product and AI teams while maintaining deep ownership of backend architecture and infrastructure.

  • Help define the operational and engineering standards for a rapidly growing platform.

  • Join at a stage where core systems, technical strategy, and engineering culture are still highly shapeable.

  • Grow into broader backend, infrastructure, or technical leadership as the team and product expand.

Please be aware that all official communication from us will come exclusively from email addresses ending in @commure.com. Any emails from other domains are not affiliated with our organization.


Employees will act in accordance with the organization’s information security policies, to include but not limited to protecting assets from unauthorized access, disclosure, modification, destruction or interference nor execute particular security processes or activities. Employees will report to the information security office any confirmed or potential events or other risks to the organization. Employees will be required to attest to these requirements upon hire and on an annual basis.

Skills Required

  • 8+ years of professional backend or infrastructure engineering experience
  • Experience designing, building, and operating production distributed systems
  • Strong proficiency in at least one modern backend programming language
  • Experience with asynchronous processing, message queues, event-driven systems, or durable workflow execution
  • Strong understanding of distributed-systems concepts including idempotency, consistency, concurrency, backpressure, retries, failure isolation, and eventual completion
  • Experience operating services in a cloud environment, including deployment, monitoring, scaling, and incident response
  • Experience designing APIs, service boundaries, and data models for complex product workflows
  • Track record of improving system reliability, observability, scalability, or operational efficiency
  • Strong debugging skills across application, infrastructure, data, and external dependency boundaries
  • Ability to reason about real-time and long-running workloads with different latency and durability requirements
  • Comfort working across backend, infrastructure, product, and AI systems
  • Ownership mindset and bias for action
  • Experience building speech-to-text, diarization, transcription, or other audio-based machine-learning workflows
  • Experience orchestrating large language models or multi-model inference pipelines in production
  • Familiarity with Temporal or similar durable execution platforms
  • Experience with Kafka, Pub/Sub, SQS, or similar event-streaming and messaging systems
  • Experience with model routing, inference gateways, rate limiting, batching, caching, or GPU-backed workloads
  • Experience building internal tooling for workflow inspection, replay, evaluation, or model debugging
  • Experience defining SLOs, managing error budgets, designing alerts, and leading production incident response
  • Familiarity with distributed tracing and observability platforms such as Grafana, OpenTelemetry, or Sentry
  • Experience with containers, Kubernetes, infrastructure as code, and cloud-native deployment systems
  • Experience optimizing systems for throughput, tail latency, infrastructure cost, and workload isolation
  • Experience with offline clients, background synchronization, or delayed and duplicated event reconciliation
  • Exposure to healthcare, HIPAA, SOC 2, encryption, privacy, security, or other regulated environments
  • Experience supporting AI-native, real-time, or agentic product experiences

Commure Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Commure and has not been reviewed or approved by Commure.

  • Fair & Transparent Compensation — Pay is considered generally market‑aligned for many roles, with engineering and senior IC/manager ranges consistent with venture‑backed health tech and major metros. Sales packages also show competitive base and OTE structures for SDRs and AEs.
  • Healthcare Strength — Core medical, dental, and vision coverage is offered, complemented by access to One Medical for convenient primary care. These elements position the health offering as robust for a mid‑size tech employer.
  • Leave & Time Off Breadth — Flexible/unlimited PTO with sick time and company holidays is provided. Parental leave is included, expanding time‑off options for different life events.

Commure Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
159 Employees
Year Founded: 2017

What We Do

Healthcare modernization doesn’t require a silver bullet looking to disrupt, it needs 1,000+ innovative solutions working together. Commure is mending fragmentation by uniting innovators across the health ecosystem to transform care with consumer-centric, data-driven digital and physical health at scale. With our universal platform and common architecture, we’re on a path to enable a system of health assurance that keeps people well while bending costs. Join us, and replace disruption with hyper-connected innovation: visit www.commure.com/careers. Commure was hatched at General Catalyst, which has backed healthcare companies such as Livongo, Oscar, Mindstrong, and Color.

Similar Jobs

Vast Logo Vast

Senior Manager, Fluid Systems Design

3D Printing • Aerospace • Hardware • Software • Manufacturing
In-Office
Long Beach, CA, USA
655 Employees
164K-233K Annually
Hybrid
San Francisco, CA, USA
289097 Employees

Enverus Logo Enverus

Account Director

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
2 Locations
1800 Employees
140K-175K Annually

Dynatrace Logo Dynatrace

Product Led Growth Marketing Lead

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
United States
5600 Employees
148K-185K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account