Staff Observability Platform Engineer (SRE)

Reposted 7 Days Ago
Be an Early Applicant
3 Locations
In-Office
118K-237K Annually
Senior level
Fitness • Healthtech • Retail • Pharmaceutical
The Role
The role focuses on designing metrics and observability frameworks, managing error budgets, and automating quality gates for release engineering, ensuring scalable cloud infrastructure and incident management.
Summary Generated by Built In

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.

POSITION SUMMARY

CVS Health PBM is looking for hands-on, passionate people who want to join a high energy and growing team, who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world. 

As a Lead Platform Reliability Engineer, you will design and implement metrics and observability frameworks with a strong focus on service level objectives (SLOs), service level indicators (SLIs), error budgets, and cloud infrastructure scaling and capacity estimation.

This individual contributor role is critical to enhancing our monitoring and observability capabilities, while also driving automation initiatives related to quality gates within the release engineering process. You will work closely with crossfunctional teams to ensure the reliability, performance, and scalable growth of our cloudbased systems.

  

Expectations for the Role:

Metrics Development: Define, implement, and maintain key performance metrics, SLOs, and SLIs to measure system reliability and performance. Ensure alignment with business objectives and operational goals.

Error Budgets: Manage error budgets effectively, collaborating with development teams to balance reliability and feature delivery. Analyze incidents and outages to inform adjustments to error budgets.

Monitoring & Observability: Design and implement comprehensive monitoring solutions to provide real-time visibility into system health. Utilize tools such as Prometheus, Grafana, Loki, Temp and other observability platforms to create dashboards and alerts.

Cloud Infrastructure Scaling: Architect, design, and implement scalable cloud infrastructure capable of supporting multiple business applications, ensuring reliability, performance, and future growth.

Quality Gates Automation: Develop and implement automated quality gates that ensure all releases meet defined reliability and performance standards. Lead the release Devops team to integrate these gates into the CI/CD pipeline.

Incident Management: Assist in incident response efforts by providing insights from metrics and monitoring tools. Conduct post-mortem analyses to identify root causes and recommend preventive measures.

AIOps Insight Automation: Use AI to surface what changed / what’s abnormal / next best action from metrics, logs, and traces—minimizing manual dashboard analysis. 

AI‑Accelerated Incident Response: Apply GenAI to speed triage and RCA with fast signal summarization and guided investigation paths. 

AI/LLM Observability & Governance: Monitor AI workloads for quality, safety, cost, latency, reliability with end‑to‑end tracing (request → prompt → tools → output) and secure logging/redaction.

AI‑Backed Release Quality Gates: Embed AI signal checks into CI/CD to flag SLO risk, latency/error drift, and regression patterns before production release.

 REQUIRED QUALIFICATIONS

  • 10+ years of experience in Software Engineering, Platform Engineering, or SRE.
  • 7+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management.
  • 7+ years building production-grade backend services in Java/python.
  • 7+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns.
  • 7+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD).
  • 7+ years working with public cloud platforms (AWS, GCP, or Azure).
  • 5+ years designing and scaling distributed, highvolume data pipelines.
  • 5+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Prometheus).
  • 5+ years with relational databases (PostgreSQL, MySQL).

 

PREFERRED QUALIFICATIONS

  • Excellent analytical skills and the ability to communicate complex technical concepts to non-technical stakeholders
  • Experience with service meshes and networking technologies such as Envoy and Istio
  • Experience integrating or operating commercial observability platforms (Splunk, AppDynamics, etc.)
  • Experience with streaming and data platforms such as Kafka, Pulsar, or similar technologies
  • Familiarity with time-series, NoSQL, or analytical databases (ClickHouse, Bigtable, Cassandra, etc.)
  • Experience with Infrastructure as Code tools such as Terraform or CloudFormation
  • Experience with cost optimization and capacity planning for large-scale cloud infra
  • Experience with chaos engineering, resiliency testing, or fault injection
  • Background in securityaware platform design, including secure servicetoservice communication
  • Experience mentoring senior engineers and influencing platform standards across organizations
  • Strong operational experience supporting 24x7 production systems, including oncall responsibilities
  • Knowledge of security best practices in cloud environments

EDUCATION

Bachelor’s degree or equivalent experience (HS diploma + 4 years relevant experience)

Pay Range

The typical pay range for this role is:

$118,450.00 - $236,900.00


This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls.  The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors.  This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above.  This position also includes an award target in the company’s equity award program. 
 

Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.

Great benefits for great people

We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.

This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility.


Additional details about available benefits are provided during the application process and on
Benefits Moments.

We anticipate the application window for this opening will close on: 08/31/2026

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.

Skills Required

  • 10+ years of experience in Software Engineering, Platform Engineering, or SRE
  • 7+ years with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management
  • 7+ years building production-grade backend services in Java/Python
  • 7+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns
  • 7+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD)
  • 7+ years working with public cloud platforms (AWS, GCP, or Azure)
  • 5+ years designing and scaling distributed, high-volume data pipelines
  • 5+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Prometheus)
  • 5+ years with relational databases (PostgreSQL, MySQL)

CVS Health Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about CVS Health and has not been reviewed or approved by CVS Health.

  • Healthcare Strength Health coverage includes medical, dental, and vision with HSA-eligible options, free preventive care, virtual care, and access to MinuteClinic services. Mental-health resources such as counseling support are emphasized, and coverage is often considered solid for full-time colleagues.
  • Retirement Support A dollar-for-dollar 401(k) match up to 5% after one year and an Employee Stock Purchase Plan are consistently highlighted in Total Rewards materials. Feedback suggests retirement programs are a meaningful strength within the overall package.
  • Wellbeing & Lifestyle Benefits Wellbeing offerings include up to 20 no-cost counseling sessions per issue, backup care, tuition assistance, and substantial in-store discounts, alongside broader wellness tools. These everyday perks expand value beyond base pay and can be especially meaningful for full-time schedules.

CVS Health Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Woonsocket, RI
119,959 Employees
Year Founded: 1963

What We Do

CVS Health is the leading health solutions company that delivers care in ways no one else can. We reach people in more ways and improve the health of communities across America through our local presence, digital channels and our nearly 300,000 dedicated colleagues – including more than 40,000 physicians, pharmacists, nurses and nurse practitioners. Wherever and whenever people need us, we help them with their health – whether that’s managing chronic diseases, staying compliant with their medications, or accessing affordable health and wellness services in the most convenient ways. We help people navigate the health care system – and their personal health care – by improving access, lowering costs and being a trusted partner for every meaningful moment of health. And we do it all with heart, each and every day.

Similar Jobs

Benchling Logo Benchling

Artificial Intelligence Engineer

Cloud • Healthtech • Social Impact • Software • Biotech
Remote or Hybrid
US
605 Employees
176K-265K Annually

Coursera + Udemy  Logo Coursera + Udemy

Director, FP&A Systems and Transformation

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
United States
1500 Employees
178K-243K Annually

PNC Bank Logo PNC Bank

Business Systems Analyst

Machine Learning • Payments • Security • Software • Financial Services
Hybrid
Farmers Branch, TX, USA
55000 Employees

PNC Bank Logo PNC Bank

Technology Solution Center Analyst

Machine Learning • Payments • Security • Software • Financial Services
Remote or Hybrid
USA
55000 Employees
37K-75K Annually

Similar Companies Hiring

Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account