Lead Infrastructure Engineer

Posted 15 Hours Ago
Be an Early Applicant
San Francisco, CA, USA
Hybrid
200K-275K Annually
Expert/Leader
Healthtech
The Role
Own Onos Health’s production infrastructure, including AWS architecture, multi-region disaster recovery, backups, monitoring, SLOs, incident response, SOC 2 and HIPAA controls, Terraform, CI/CD, preview environments, and AI-agent deployment guardrails. Establish platform strategy, reliability practices, security automation, audit evidence, and on-call operations as the company’s first dedicated infrastructure hire. The role also requires technical leadership, delegation, and communication with non-technical stakeholders.
Summary Generated by Built In
About Onos Health

Onos Health’s mission is simple but ambitious: ensure every healthcare dollar goes toward delivering the highest quality care. Today, 30% of total U.S. healthcare spending is wasted due to ineffective care and administrative burden caused by misalignment between providers and payers.

Onos is addressing this by building the largest AI-driven healthcare data platform. Our models enables payers to make faster, more accurate decisions across their populations. By guiding members to the right care, Onos is channeling more dollars to high-quality care that drives better outcomes while making healthcare more affordable.

Onos is well-funded by some of the best healthcare investors and is working with the nation’s largest health plans. Come join a category-defining company and help reimagine healthcare for the better.

Why Onos?
  • Meaningful impact: Help fix what is fundamentally broken in healthcare

  • Direct collaboration: Work alongside experienced founders with deep healthcare and data expertise

  • Culture: Join a high-performing, transparent, and results-oriented team

  • Ownership: Significant responsibility and autonomy from day one

  • Opportunity: Play a pivotal role in building a fast-growing, category-defining healthcare AI company

The Role

We're seeking an experienced infrastructure engineer to become our first dedicated platform hire and the owner of the infrastructure the Onos platform runs on. Onos is in production with the nation's largest health plans, which comes with contractual uptime SLAs, disaster recovery commitments, and a security bar (SOC 2, HIPAA) our clients audit. Until now this has been carried collectively by our product engineers and founders — you'll own it end to end. We build heavily with AI coding agents, so much of your leverage will come from specifying work well and directing agents rather than typing every line yourself; prior tech lead or engineering management experience translates directly. As an early team member, you'll set the patterns every future platform engineer at Onos inherits. This role is a hybrid role based in San Francisco, where you'll be expected to work at our office in person 3 times a week.

What you'll be doing at Onos:
  • Own our availability, disaster recovery, and backup commitments to enterprise clients — multi-region failover architecture and the recovery exercises that prove it

  • Stand up production monitoring, alerting, SLOs, and our on-call rotation and incident response process

  • Own the technical controls behind SOC 2 and HIPAA: cloud security posture (AWS org guardrails, IAM least-privilege, KMS/encryption), vulnerability remediation, and continuous audit evidence through Vanta, with a path toward HITRUST

  • Build the CI/CD pipelines, Terraform/IaC foundations, preview environments, and test infrastructure the whole team ships on

  • Build the guardrails that let AI coding agents ship safely — policy-as-code, deploy verification, and agent-operated operations tooling

  • Set the strategy and operating rhythm for platform work: priorities, status, and what we deliberately defer

Technical Challenges At Onos:
  • Architect AI SRE agents to ensure up-to-date compliance and reliability, enabling engineers to work more effectively and strategically

  • Right-size enterprise-grade reliability: SLOs and alerting you can trust without drowning a small team in pager noise

  • Turn compliance into continuously verified infrastructure — security controls and audit evidence as code

  • Scale a CI/CD and environments platform where AI agents, not just humans, are the primary users

  • Support multi-region disaster recovery with defined RTO/RPO targets and immutable, restore-tested backups for a multi-tenant healthcare platform

Tech Stack:

At Onos, we work with a modern tech stack where we continuously evaluate and adopt cutting-edge technologies as we scale.

  • Infrastructure/Systems: AWS (ECS, Bedrock, Glue, etc.), Langfuse, Terraform

  • Languages/Frameworks

    • Backend: Python, Django, Celery / Celery Beat, django-ninja, django-tenants

    • Frontend: NextJS, Typescript, Tanstack Query, Shadcn UI, Zod, Nuqs

  • Database/Storage: PostgreSQL (AWS RDS), S3, Clickhouse

  • Development Tools: Github, Linear, Claude Code, Codex, CoderabbitAI

What we're looking for:
  • Deep AWS experience — you've owned production cloud infrastructure end to end (IAM, networking, KMS, containers, managed databases)

  • Strong Terraform/IaC and CI/CD expertise; you treat pipelines and environments as products with users

  • Taken a company through at least one SOC 2 (or HITRUST/ISO 27001) audit cycle with your hands on the technical controls

  • SRE fundamentals — SLOs, incident management, DR design — with the judgment to right-size reliability for a startup with enterprise contracts

  • Prior experience leading engineering teams as a tech lead or engineering manager — you can break down ambiguous goals, delegate (to humans or AI agents), and communicate crisply with non-technical stakeholders

  • Customer obsessed and motivated to make an impact in the healthcare space

Bonus points if you have:
  • Significant experience in healthcare or another regulated industry, with HIPAA fluency

  • Policy-as-code (OPA, Kyverno) or compliance automation (e.g., Vanta, Drata) experience

  • Been the first infrastructure hire, or helped found or lead a platform team

  • Built internal tooling or infrastructure for LLM/agent systems

Benefits and Perks
  • Hybrid arrangement: 3 days/week at San Francisco office (Financial District)

  • Unlimited vacation policy

  • Paid parental leave

  • Medical, dental, and vision insurance

  • Pre-tax commuter benefits

  • 401(k)

  • Significant equity as an early employee

  • Direct mentorship from experienced founders

  • Ground-floor opportunity to help build a team and culture

  • Regular team events and offsites

  • Company-provided equipment and home office setup

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Skills Required

  • Deep AWS experience owning production cloud infrastructure, including IAM, networking, KMS, containers, and managed databases
  • Strong Terraform or infrastructure-as-code expertise
  • Strong CI/CD expertise, including pipelines and environment platforms
  • Hands-on experience taking a company through at least one SOC 2, HITRUST, or ISO 27001 audit cycle
  • Experience with SRE fundamentals, including SLOs, incident management, and disaster recovery design
  • Prior experience leading engineering teams as a technical lead or engineering manager
  • Ability to break down ambiguous goals, delegate work to humans or AI agents, and communicate with non-technical stakeholders
  • Customer-focused mindset and motivation to work in healthcare
  • Significant healthcare or regulated-industry experience and HIPAA fluency
  • Experience with policy-as-code tools such as OPA or Kyverno
  • Experience with compliance automation tools such as Vanta or Drata
  • Experience as a first infrastructure hire or in founding or leading a platform team
  • Experience building internal tooling or infrastructure for LLM or agent systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
8 Employees

What We Do

There are too many reasons why Healthcare is unaffordable in the U.S, but one of the largest is waste. Up to 40% of healthcare spend is lost on inefficiencies, administrative costs, and ineffective care practices. We’re building a system that unites providers and payers with one goal: delivering the highest quality care while eliminating waste.

Similar Jobs

Hybrid
San Francisco, CA, USA
289097 Employees

Wells Fargo Logo Wells Fargo

Infrastructure Engineer

Fintech • Financial Services
Hybrid
San Francisco, CA, USA
205000 Employees
119K-224K Annually
Hybrid
2 Locations
289097 Employees
Hybrid
2 Locations
289097 Employees

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account