Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Gurgaon, Gurugram, Haryana, IND
In-Office
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
The Role
Lead site reliability engineering across system availability, performance, scalability, observability, incident response, automation, CI/CD, and cloud migrations. Architect OpenTelemetry-based monitoring solutions, improve MTTD and MTTR, develop infrastructure tooling, support on-premises-to-cloud migrations, and apply AI-driven operations. Collaborate with globally distributed engineering teams on performance testing, troubleshooting, postmortems, and continuous reliability improvements.
Summary Generated by Built In
Requisition Number: 2385995
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
We are looking for a highly skilled Sr. Site Reliability Engineer (SRE) to join our India. As a Sr. SRE, you will be responsible for ensuring the reliability, performance, and scalability of our systems. You will contribute to key projects, including performance testing, CI/CD tooling, and facilitating infrastructure and application migrations while working closely with both the India team and the existing SRE team in the United States.
Primary Responsibilities:
  • System Reliability: Ensure the availability, performance, and scalability of critical systems by implementing best practices in site reliability engineering
  • Observability & Telemetry: Drive the design and evolution of observability systems by building scalable, extensible solutions using OpenTelemetry (OTEL) and other modern observability tools. Champion innovation in monitoring, distributed tracing, and logging strategies to provide deep visibility into system behavior. Continuously evaluate and integrate emerging technologies to improve observability maturity and reduce mean time to detect (MTTD) and resolve (MTTR)
  • Project Leadership: Lead and contribute to projects such as performance testing, CI/CD tooling, and infrastructure/application migrations with focus to migrate from on-prem to cloud solutions
  • Incident Response: Actively participate in incident response, troubleshooting, and post-mortem analysis to identify root causes and prevent future occurrences
  • Automation and Tooling: Develop and maintain automation tools to reduce manual effort, streamline processes, and enhance system reliability
  • Collaboration: Work closely with other SREs, engineers, and stakeholders across time zones to align on goals, strategies, and ensure smooth project execution
  • Continuous Improvement: Identify opportunities to improve system reliability, performance, and operational efficiency, and implement changes as needed
  • AI Driven Operations: Leverage AI-powered tools and platforms to enhance observability, incident response, and operational efficiency
  • Comply with all applicable Company policies, procedures, and business directives, changes including those relating to work location, team assignments, work schedules, and flexible work arrangements

Required Qualifications:
  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)
  • 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role
  • Observability Expertise: Proven experience architecting and implementing observability platforms using OpenTelemetry, Datadog, Splunk, Grafana, or similar tools. Demonstrated ability to innovate in this space - whether by building custom telemetry pipelines, integrating AI/ML for anomaly detection, or developing new approaches to visualize and interpret system health
  • Containers & Orchestration: Experience with Docker and Kubernetes
  • CI/CD: Experience with CI/CD tools like Jenkins, GitHub Actions, and related automation pipelines
  • Cloud Platforms: Solid knowledge of public cloud platforms, preferrably Azure, and expertise in On-Prem to Cloud migrations
  • Technical Expertise: Deep understanding of systems architecture, cloud infrastructure, networking, and automation tools
  • AI Tools: Proven exposure to AI tools and their application in SRE workflows for faster delivery and smarter operations
  • Automation Skills: Proven solid scripting/programming skills (Python, Go, Powershell, Bash, etc), and experience with infrastructure-as-code tools like Terraform and Ansible
  • Problem-Solving: Proven excellent problem-solving skills, with experience in incident management, troubleshooting, and root cause analysis
  • Collaboration: Proven excellent communication and collaboration skills, with the ability to work effectively in a distributed team across time zones

Preferred Qualifications:
  • Experience working in a global or distributed team environment
  • Industry experience in Payments, Fintech, or Healthcare
  • Knowledge of security best practices in cloud and distributed systems

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience
  • 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role
  • Experience architecting and implementing observability platforms using OpenTelemetry, Datadog, Splunk, Grafana, or similar tools
  • Experience with Docker and Kubernetes
  • Experience with CI/CD tools such as Jenkins and GitHub Actions
  • Knowledge of public cloud platforms, preferably Azure
  • Experience with on-premises-to-cloud migrations
  • Deep understanding of systems architecture, cloud infrastructure, networking, and automation tools
  • Exposure to AI tools and their application in SRE workflows
  • Solid scripting or programming skills in Python, Go, PowerShell, Bash, or similar languages
  • Experience with infrastructure-as-code tools such as Terraform and Ansible
  • Experience in incident management, troubleshooting, and root cause analysis
  • Excellent communication and collaboration skills across distributed teams and time zones
  • Experience working in a global or distributed team environment
  • Industry experience in Payments, Fintech, or Healthcare
  • Knowledge of security best practices in cloud and distributed systems

What the Team is Saying

Optum Compensation & Benefits Highlights

  • Leave & Time Off Breadth — PTO is generally described as decent or good, and many note it as a strong part of the package. Actual ability to take time off can depend on workload and team coverage.
  • Retirement Support — Offerings include a 401(k) with employer match and access to an employee stock purchase plan, which are highlighted as meaningful components of total rewards. These programs are consistently referenced among core benefits.
  • Parental & Family Support — Parental and caregiver leave, along with adoption assistance, are publicly highlighted and viewed as notable elements of the package. Availability can be role-specific, but these supports contribute to overall breadth where offered.

Optum Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Eden Prairie, MN
160,000 Employees
Year Founded: 2011

What We Do

Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.

Gallery

Gallery
Gallery
Gallery

Optum Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.

Typical time on-site: Not Specified
HQEden Prairie, MN
Metro Manila, Philippines
Cebu, Philippines
Davao, Philippines
Ann Arbor, MI
Atlanta, GA
Baltimore, MD
Bengaluru, India
Chennai, India
Dallas, TX
Detroit, MI
Dublin, Ireland
Hartford, CT
Houston, TX
Hyderabad, India
Jacksonville, FL
Las Vegas, NV
Letterkenny, Ireland
Louisville, KY
Madison, WI
Minneapolis, MN
Nashville, TN
New Delhi, India
Philadelphia, PA
Phoenix, AZ
Pune, India
Raleigh, NC
San Diego, CA
Washington, DC
Learn more

Similar Jobs

Optum Logo Optum

Senior Software Engineering Lead - Java and Azure

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Gurgaon, Gurugram, Haryana, IND
160000 Employees

Optum Logo Optum

Designer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Gurgaon, Gurugram, Haryana, IND
160000 Employees

Optum Logo Optum

Senior Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Gurgaon, Gurugram, Haryana, IND
160000 Employees

Optum Logo Optum

Manager Data Science - AI ML and Snowflake

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Gurgaon, Gurugram, Haryana, IND
160000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account