Site Reliability Engineer III

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
175K-185K Annually
Senior level
Mobile
The Role
Serve as Vida’s first dedicated Site Reliability Engineer, modernizing Terraform, standardizing environments, improving CI/CD, scaling GCP and Kubernetes infrastructure, upgrading databases and runtimes, and strengthening monitoring, observability, security, and operational processes. The role includes building runbooks, on-call and escalation procedures, supporting enterprise launches, managing infrastructure costs, and evaluating multi-cluster Kubernetes architecture in a fully remote environment.
Summary Generated by Built In
ABOUT US
 
At Vida, we help people get better- and we're helping the healthcare system get better, too.
 
Vida is a virtual, personalized obesity care provider that uses evidence-based treatment to help patients manage obesity and related conditions like diabetes, high blood pressure, anxiety and depression. Vida's team of Obesity Medicine-Certified Physicians, Registered Dietitians, Expert Coaches and Licensed Therapists takes a whole-person approach to care, helping people lose weight, reduce stress and improve their overall health.
 
By combining advanced technology with top-notch healthcare providers, Vida is breaking down the barriers that have historically kept people from getting the best care. It's trusted by Fortune 100 companies, major national payers and large providers to enable their employees to live their healthiest lives.

Vida has been operating and growing for years, and our infrastructure reflects that. We run on GCP with a production GKE cluster hosting around 50 workloads, from Django applications to scheduled Airflow jobs. Our data layer includes Cloud SQL (MySQL and PostgreSQL), Redis, and Firestore. Our infrastructure is defined in two Terraform repositories, one for core GCP infrastructure and one for our data platform, and both have grown across many contributors over time. Until now, our infrastructure has been managed by backend engineers with deep infrastructure experience, and this role adds our first dedicated SRE to that group.

You'll be Vida's first dedicated Site Reliability Engineer. You'll join the Enablement Team, which owns the platform and tooling our Engineering Teams build on. You'll report to the Engineering Manager and work closely with the team's Lead Engineer, who sets technical direction and will mentor you. This is a fully remote role with no time zone restrictions.

You'll modernize, consolidate, and scale our infrastructure as Vida takes on a wave of new enterprise contracts starting January 1. You'll also help shape what SRE looks like at Vida going forward.

Repsonsibilities:

  • Consolidate our Terraform, which has grown into inconsistent patterns across our infrastructure and data repositories, into a clean, well-documented structure the whole team can work in. Establish conventions for state management, module structure, code review, and CI checks.
  • Normalize environments, improve build and deploy automation in GitHub Actions, and add drift detection and alerting.
  • Apply overdue patches and upgrades across our Cloud SQL databases and application runtimes.
  • Right-size compute and database workloads for growth, including connection pooling and scaling improvements for high-traffic services.
  • Evaluate our Kubernetes architecture as we grow, including whether and when to move to a multi-cluster setup.
  • Improve monitoring and observability in Datadog and Cloud Monitoring so we catch issues before they become incidents.
  • Design observability access for contractors and external partners that gives them the visibility they need while keeping protected health information out of view.
  • Retire legacy infrastructure and tooling that has been replaced but not yet decommissioned.
  • Build repeatable operational processes, including runbooks, an on-call rotation, and escalation documentation.
  • Support infrastructure readiness for Vida's January 1 enterprise launches.
  • Additional responsibilities as needed.

Qualifications:

  • Bachelor's degree at a minimum.
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
  • Deep hands-on Terraform experience, including structuring modules and managing state across environments.
  • Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
  • Production Kubernetes experience, including autoscaling, resource management, and judgment about what belongs in the cluster versus outside it.
  • Hands-on experience building monitoring, alerting, and dashboards with tools like Datadog or Cloud Monitoring.
  • Proficiency in Python for tooling and automation.
  • Comfortable working across multiple teams and disciplines, and explaining infrastructure decisions to non-specialists.

Preferred:

  • Experience as an early or first SRE hire.
  • Experience refactoring or consolidating a large, organically grown Terraform codebase.
  • Experience improving observability from a less mature baseline.
  • Experience in a HIPAA-regulated or other compliance-driven environment.
  • CI/CD experience with GitHub Actions.
  • Experience running Django applications or Airflow in production on Kubernetes.
  • Experience designing or migrating to multi-cluster Kubernetes architectures.

Vida is proud to be an Equal Employment Opportunity and Affirmative Action employer.
 
Diversity is more than a commitment at Vida—it is the foundation of what we do. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, gender, gender identity or expression, sexual orientation, marital status, national origin, genetics, disability, age, or Veteran status. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law.
 
We seek to recruit, develop and retain the most talented people from a diverse candidate pool. We don’t just accept differences — we celebrate them, we support them, and we thrive on them for the benefit of our employees, our platform and those we serve. Vida is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.
 
We do not accept unsolicited assistance from any headhunters or recruitment firms for any of our job openings. All resumes or profiles submitted by search firms to any employee at Vida in any form without a valid, signed search agreement in place for the specific position will be deemed the sole property of Vida. No fee will be paid in the event the candidate is hired by Vida as a result of the unsolicited referral.
 
**Vida is authorized to do business in many, but not all, states. If you are not located in or able to work from a state where Vida is registered, you will not be eligible for employment. Please speak with your recruiter to learn more about where Vida is registered.
 
Please note: Applicants must be authorized to work in the U.S. as Vida is unable to sponsor work visas for any position.
All Vida Employees must reside in/be able to work from the U.S.- international work is prohibited. Job postings at Vida will remain open through end of year, until filled.
 
#LI-remote

Skills Required

  • Bachelor's degree
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering
  • Production systems ownership experience
  • Deep hands-on Terraform experience, including module structure and state management across environments
  • Strong working knowledge of GCP, including GKE, Cloud SQL, IAM, networking, load balancing, and cost management
  • Production Kubernetes experience, including autoscaling and resource management
  • Experience building monitoring, alerting, and dashboards with Datadog or Cloud Monitoring
  • Proficiency in Python for tooling and automation
  • Ability to work across multiple teams and explain infrastructure decisions to non-specialists
  • Experience as an early or first SRE hire
  • Experience refactoring or consolidating a large Terraform codebase
  • Experience improving observability from a less mature baseline
  • Experience in a HIPAA-regulated or compliance-driven environment
  • CI/CD experience with GitHub Actions
  • Experience running Django applications or Airflow in production on Kubernetes
  • Experience designing or migrating to multi-cluster Kubernetes architectures
  • Authorization to work in the United States without visa sponsorship
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
636 Employees
Year Founded: 2014

What We Do

Vida is a virtual care company that combines a human-centric approach with technology to address chronic and co-occurring physical and behavioral health conditions. We provide personalized chronic condition management combined with health coaching and therapy through a mobile and online platform that supports individuals in managing and significantly improving conditions such as diabetes, hypertension, obesity, depression, anxiety, etc. Our platform integrates deeply individual expert care with machine learning and remote monitoring to deliver lasting behavior change, health outcomes and cost savings. Vida is in the business of enabling self-insured employers, health plans and providers to take better care of their employees and members. We are trusted by Fortune 1000 companies, major national payers, and large providers to activate, engage, and empower their employees to live their healthiest lives. Based in San Francisco, CA, Vida is backed by investors including Khosla Ventures, StartX, Aspect Ventures, Canvas, Workday, and Nokia.

Similar Jobs

Backblaze Logo Backblaze

Site Reliability Engineer

Cloud • Information Technology
Remote
United States
363 Employees
125K-150K Annually

onXmaps, Inc. Logo onXmaps, Inc.

Site Reliability Engineer

Consumer Web • Information Technology • Mobile • Other • Software • App development
In-Office or Remote
9 Locations
450 Employees
130K-153K Annually

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
10 Locations
5550 Employees
127K-249K Annually

Similar Companies Hiring

Prolaio Thumbnail
Artificial Intelligence • Big Data • Healthtech • Mobile • Wearables • Analytics
Chicago, IL
82 Employees
ARB Interactive Thumbnail
Gaming • Mobile • Software
Miami, Florida
190 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account