Senior Engineer - Observability Platform

Posted Yesterday
Be an Early Applicant
27 Locations
In-Office or Remote
83K-222K Annually
Senior level
Fitness • Healthtech • Retail • Pharmaceutical
The Role
Design, build, and operate large-scale observability services and telemetry pipelines using Go, Python, and Java. Lead enterprise adoption of OpenTelemetry, implement monitoring/SLOs, automate platform operations, participate in on-call rotations, and collaborate with SRE, Cloud, and security teams to ensure reliable, secure, and cost-efficient observability at scale.
Summary Generated by Built In

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.

POSITION SUMMARY

Join CVS Health Enterprise Technology and help evolve observability at Fortune‑6 scale. The Enterprise Observability Platform (EOP) delivers standardized, frictionless instrumentation and telemetry pipelines for engineering teams across all CVS Health application environments—spanning on‑prem, hybrid, and multiple public clouds.

As a Senior Observability Platform Engineer, you will design, build, and operate large‑scale observability services that process billions of logs, metrics, and traces daily. You will develop high‑performance backend services using Go, Python, and Java, and lead the adoption of OpenTelemetry-based instrumentation and standards across the enterprise.

In this role, you will partner closely with SRE, Cloud Engineering, CI/CD, Infrastructure, Digital Security, and application teams to shape platform strategy, enhance developer experience, and ensure reliable, secure, and cost‑efficient observability at scale. You will provide senior technical leadership, influence architectural direction, and help deliver a world‑class, self-service observability ecosystem that accelerates engineering productivity and operational excellence.

Responsibilities for this role

  • Develop custom software to drive the observability platform using technologies such as Java Spring Boot, Python, Go-lang etc
  • Design, build, and operate core observability platform services using Go, Python, Java (Spring Boot)
  • Support and lead enterprise-wide adoption of OpenTelemetry, including client libraries, semantic conventions, instrumentation patterns, and Collector/agent strategy
  • Design, architect and scale high‑throughput, fault‑tolerant telemetry pipelines (logs, metrics, traces) with a focus on performance, reliability, and cost efficiency
  • Develop self-service observability capabilities that simplify onboarding, troubleshooting, and adoption for application teams
  • Implement end-to-end monitoring of the observability platform itself, defining SLOs, health checks, and alerting
  • Collaborate with SRE, Platform, and Cloud teams to establish reliability standards, error budgets, and incident response practices
  • Participate in on‑call rotations and lead incident mitigation, root‑cause analysis, and post‑incident reviews
  • Automate operational workflows and eliminate manual toil through tooling, CI/CD enhancements, and platform automation
  • Ensure secure telemetry pipelines through mTLS, secrets management, and zero‑trust design patterns
  • Produce and maintain high-quality technical documentation, standards, and best practices
  • Engage with internal engineering teams to gather requirements, influence roadmap prioritization, and deliver platform improvements
  • Provide technical leadership through mentorship, design reviews, architectural guidance, and cross‑team collaboration with principal engineers and engineering leadership

REQUIRED QUALIFICATIONS

  • 5+ years of experience in Software Engineering
  • 3+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management
  • 3+ years building production-grade backend services in Go and/or Java
  • 3+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns
  • 3+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD)
  • 3+ years working with public cloud platforms (AWS, GCP, or Azure)
  • 2+ years designing and scaling distributed, high‑volume data pipelines
  • 2+ years of experience with Infrastructure as Code tools such as Terraform or CloudFormation
  • 2+ years of experience with Helmcharts, Kustomize etc
  • 2+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Mimir)
  • 2+ years with relational databases (PostgreSQL, MySQL)

PREFERRED QUALIFICATIONS

  • Experience with service meshes and networking technologies such as Envoy and Istio
  • Experience integrating or operating commercial observability platforms (Datadog, New Relic, AppDynamics, etc.)
  • Experience with On-Call Scheduling tools such OpsGenie, PagerDuty, GoAlert etc.
  • Experience with streaming and data platforms such as Kafka, Pulsar, or similar technologies
  • Familiarity with time-series, NoSQL, or analytical databases (ClickHouse, Bigtable, Cassandra, etc.)
  • Experience with cost optimization and capacity planning for large-scale telemetry systems
  • Experience with chaos engineering, resiliency testing, or fault injection
  • Background in security‑aware platform design, including secure service‑to‑service communication
  • Experience mentoring senior engineers and influencing platform standards across organizations
  • Strong operational experience supporting 24x7 production systems, including on‑call responsibilities
  • Strong technical communication and cross‑team collaboration skills
  • Experience operating in regulated or compliance‑heavy environments (e.g., healthcare, finance)

EDUCATION

Bachelor’s degree from accredited university or equivalent work experience (HS diploma + 4 years relevant experience)

BUSINESS OVERVIEW
Bring your heart to CVS Health Every one of us at CVS Health shares a single, clear purpose: Bringing our heart to every moment of your health. This purpose guides our commitment to deliver enhanced human-centric health care for a rapidly changing world. Anchored in our brand — with heart at its center - our purpose sends a personal message that how we deliver our services is just as important as what we deliver.  Our Heart At Work Behaviors™ support this purpose. We want everyone who works at CVS Health to feel empowered by the role they play in transforming our culture and accelerating our ability to innovate and deliver solutions to make health care more personal, convenient and affordable.  We strive to promote and sustain a culture of diversity, inclusion and belonging every day.  CVS Health is an affirmative action employer, and is an equal opportunity employer, as are the physician-owned businesses for which CVS Health provides management services. We do not discriminate in recruiting, hiring, promotion, or any other personnel action based on race, ethnicity, color, national origin, sex/gender, sexual orientation, gender identity or expression, religion, age, disability, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.  We proudly support and encourage people with military experience (active, veterans, reservists and National Guard) as well as military spouses to apply for CVS Health job opportunities

Anticipated Weekly Hours

40

Time Type

Full time

Pay Range

The typical pay range for this role is:

$83,430.00 - $222,480.00

This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls.  The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors.  This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above. 
 

Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.

Great benefits for great people

We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.

This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility.


Additional details about available benefits are provided during the application process and on
Benefits Moments.

We anticipate the application window for this opening will close on: 10/30/2026

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.

Skills Required

  • 5+ years of experience in Software Engineering
  • 3+ years of experience with observability practices (SLIs/SLOs/SLAs, alerting, incident management)
  • 3+ years building production-grade backend services in Go and/or Java
  • 3+ years implementing and operating OpenTelemetry (including OTLP, semantic conventions, instrumentation patterns)
  • 3+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD)
  • 3+ years working with public cloud platforms (AWS, GCP, or Azure)
  • 2+ years designing and scaling distributed, high-volume data pipelines
  • 2+ years of experience with Infrastructure as Code tools (Terraform or CloudFormation)
  • 2+ years of experience with Helmcharts and Kustomize
  • 2+ years working with Grafana OSS or comparable observability backends (Grafana, Loki, Tempo, Mimir)
  • 2+ years with relational databases (PostgreSQL, MySQL)
  • Bachelor's degree or equivalent work experience
  • Experience with service meshes and Envoy/Istio
  • Experience integrating or operating commercial observability platforms (Datadog, New Relic, AppDynamics)
  • Experience with on-call scheduling tools (OpsGenie, PagerDuty, GoAlert)
  • Experience with streaming/data platforms (Kafka, Pulsar)
  • Familiarity with time-series or NoSQL/analytical databases (ClickHouse, Bigtable, Cassandra)
  • Experience with cost optimization and capacity planning for telemetry systems
  • Experience with chaos engineering, resiliency testing, or fault injection
  • Background in security-aware platform design (mTLS, secrets management, zero-trust)
  • Experience mentoring senior engineers and influencing platform standards
  • Operational experience supporting 24x7 production systems and on-call responsibilities
  • Experience operating in regulated or compliance-heavy environments (healthcare, finance)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Woonsocket, RI
119,959 Employees
Year Founded: 1963

What We Do

CVS Health is the leading health solutions company that delivers care in ways no one else can. We reach people in more ways and improve the health of communities across America through our local presence, digital channels and our nearly 300,000 dedicated colleagues – including more than 40,000 physicians, pharmacists, nurses and nurse practitioners. Wherever and whenever people need us, we help them with their health – whether that’s managing chronic diseases, staying compliant with their medications, or accessing affordable health and wellness services in the most convenient ways. We help people navigate the health care system – and their personal health care – by improving access, lowering costs and being a trusted partner for every meaningful moment of health. And we do it all with heart, each and every day.

Similar Jobs

Motive Logo Motive

Lead, Safety and Compliance Strategy (Remote USA)

Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Easy Apply
Remote
United States
4000 Employees
135K-170K Annually

General Motors Logo General Motors

GM Defense Controllership Senior Analyst

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
United States
165000 Employees
88K-141K Annually

Runpod Logo Runpod

Technical Program Manger

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Remote
USA
80 Employees
140K-165K Annually

PwC Logo PwC

US Tech - AI Engineering Senior Associate

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Remote or Hybrid
68 Locations
370000 Employees
151K-187K Annually

Similar Companies Hiring

Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account