Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Chennai, Tamil Nadu, IND
In-Office
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
The Role
Lead the design, implementation, and operation of AI Ops and SRE solutions in public cloud environments. Build RAG pipelines, agentic workflows, and AI-powered enterprise applications; automate infrastructure with Terraform and GitHub Actions; manage Kubernetes; lead incident response; ensure reliability, scalability, and performance; establish AI evaluation and monitoring frameworks; and mentor engineering teams.
Summary Generated by Built In
Requisition Number: 2378612
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
  • Lead end-to-end design and implementation of AI Ops solutions from concept through production with an emphasis on responsible AI practices
  • Contribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments
  • Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practices
  • Define and own solution architecture for RAG pipelines, agentic workflows, tool calling, orchestration, and conversational context management
  • Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions
  • Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers
  • Design, code, test, and operate software using Python & Node.js
  • Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency
  • Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so

Required Qualifications:
  • Bachelor's degree in information systems, Computer Science, Engineering, or related field or equivalent certification
  • 5+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure)
  • 3+ years of demonstrated hands-on experience with Python and Terraform based development
  • 2+ years delivering AI/ML or Generative AI solutions in production
  • 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE, or self-hosted)
  • Available to work rotating 24x7 primary and secondary on-call shifts

Preferred Qualifications:
  • Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
  • Experience leading globally distributed technical teams
  • Experience with vector search and enterprise search solutions
  • CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform
  • Proven experience building Retrieval Augmented Generation (RAG) pipelines, Agentic AI, or multi-step AI workflows
  • Proven experience building evaluation and monitoring frameworks for AI quality
  • Proven effective communication skills with ability to explain complex technical concepts to diverse stakeholders

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.

Skills Required

  • Bachelor's degree in information systems, computer science, engineering, or a related field, or equivalent certification
  • 5+ years of software engineering and Site Reliability Engineering experience in a public cloud environment
  • 3+ years of hands-on Python and Terraform development experience
  • 2+ years delivering AI, machine learning, or generative AI solutions in production
  • 2+ years managing Kubernetes environments, including EKS, AKS, GKE, or self-hosted Kubernetes
  • Availability to work rotating 24x7 primary and secondary on-call shifts
  • Experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
  • Experience leading globally distributed technical teams
  • Experience with vector search and enterprise search solutions
  • CI/CD and Infrastructure as Code experience, preferably with GitHub Actions and Terraform
  • Experience building RAG pipelines, agentic AI, or multi-step AI workflows
  • Experience building evaluation and monitoring frameworks for AI quality
  • Effective communication skills and ability to explain complex technical concepts to diverse stakeholders

What the Team is Saying

Optum Compensation & Benefits Highlights

  • Healthcare Strength Official materials highlight copay and HSA medical plan choices with in‑network preventive care at 100%, prescription coverage, and low/no‑cost virtual visits, plus company HSA contributions. Dental preventive services are 100% in network, and mental health resources include an EAP and premium Calm access.
  • Parental & Family Support Programs include six weeks paid parental leave, up to two weeks paid caregiver leave, and Bright Horizons back‑up care with enhanced family supports. Adoption assistance up to $10,000 for full‑time employees reinforces family‑oriented benefits.
  • Equity Value & Accessibility Financial benefits include an Employee Stock Purchase Plan at a 10% discount, expanding access to equity ownership.

Optum Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Eden Prairie, MN
160,000 Employees
Year Founded: 2011

What We Do

Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.

Gallery

Gallery
Gallery
Gallery

Optum Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.

Typical time on-site: Not Specified
HQEden Prairie, MN
Metro Manila, Philippines
Cebu, Philippines
Davao, Philippines
Ann Arbor, MI
Atlanta, GA
Baltimore, MD
Bengaluru, India
Chennai, India
Dallas, TX
Detroit, MI
Dublin, Ireland
Hartford, CT
Houston, TX
Hyderabad, India
Jacksonville, FL
Las Vegas, NV
Letterkenny, Ireland
Louisville, KY
Madison, WI
Minneapolis, MN
Nashville, TN
New Delhi, India
Philadelphia, PA
Phoenix, AZ
Pune, India
Raleigh, NC
San Diego, CA
Washington, DC
Learn more

Similar Jobs

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Optum Logo Optum

Machine Learning Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Optum Logo Optum

Manager Collections

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account