Site Reliability Engineer

Reposted 16 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
200K-275K Annually
Senior level
Artificial Intelligence • Healthtech • Information Technology • Software
The Role
As a Site Reliability Engineer, you will manage the production environment, focusing on infrastructure design, automation, and optimizing deployment pipelines to ensure high availability.
Summary Generated by Built In
SRE

Location: San Francisco, CA (5 Days In-Office)

You are the infrastructure expert who enables our rapid product development and guarantees 99.9%+ stability and performance of our clinical AI platform for major health systems. Your focus on operational excellence is directly tied to a patient's access to life-saving treatment.

What We Look for in a Great Engineer

You have the intensity and technical mastery to own mission-critical infrastructure. You hold yourself and others to high standards and thrive in a high-energy, in-office culture where everyone is in it to win it.

  • Tool Proficiency: You are highly proficient with your tools—you speak command line fluently and have mastered keyboard shortcuts.

  • Ownership: You thrive on owning complex systems and have a proven track record of scaling mission-critical deployments.

  • Automation Drive: You love automating things, always finding new ways to increase your own leverage, and defining standards for operational excellence.

  • Problem Solver: You won't wait for someone else to solve a problem that you're in a position to solve; you are willing to jump into whatever needs to get done.

What You'll Work On (Responsibilities)

As our SRE, you will own the entire production environment and improve the development experience:

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.

  • Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.

  • CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.

  • DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.

  • Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.

Technical Qualifications & Environment

  • IaC & Orchestration: Deep, demonstrable experience with Kubernetes, Helm, and Terraform.

  • Scaling Systems: Proven ability to architect and maintain complex, distributed systems with high-availability requirements.

  • Deployment Experience: Hands-on experience optimizing deployment pipelines for both application code (TypeScript) and machine learning models (Python/ML). Also PostgreSQL, Redis, Kakfa.

  • Core Team Member: Excitement about working five days per week in our San Francisco office.

Skills Required

  • Deep experience with Kubernetes, Helm, and Terraform
  • Ability to architect and maintain complex distributed systems
  • Hands-on experience with TypeScript and Python/ML deployment pipelines
  • Proficiency with PostgreSQL, Redis, and Kafka
  • Previous experience handling 500+ machine deployments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
49 Employees
Year Founded: 2022

What We Do

Latent is redefining medication access through clinical AI, partnering with leading health systems to streamline prior authorizations, 340B compliance, and appeals. By optimizing pharmacy workflows and enabling centralization, Latent accelerates time-to-therapy for patients, supports scalable pharmacy expansion, and reduces administrative burdens on staff. Learn more about our mission at www.latenthealth.com/.

Similar Jobs

CrowdStrike Logo CrowdStrike

Site Reliability Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
Sunnyvale, CA, USA
10000 Employees
120K-180K Annually

ServiceNow Logo ServiceNow

Site Reliability Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Santa Clara, CA, USA
29000 Employees
166K-290K Annually

BAE Systems, Inc. Logo BAE Systems, Inc.

Site Reliability Engineer

Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Hybrid
San Diego, CA, USA
40000 Employees
118K-201K Annually

Sprinter Health Logo Sprinter Health

Site Reliability Engineer

Artificial Intelligence • Healthtech • Logistics • Social Impact • Software • Telehealth
Remote or Hybrid
2 Locations
500 Employees
160K-255K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account