Lead Cloud Reliability Engineer - Remote

Posted An Hour Ago
Hiring Remotely in Minnetonka, MN, USA
In-Office or Remote
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
The Role
Leads cloud reliability and SRE engineering for enterprise AI operations. Responsibilities include automating infrastructure with Terraform and Python, managing Kubernetes environments, leading incident response, designing highly available cloud systems, deploying AI and RAG solutions, improving observability and operational efficiency, and mentoring globally distributed engineering teams. The role requires rotating 24x7 on-call coverage and hands-on leadership across cloud infrastructure, automation, reliability, and responsible AI practices.
Summary Generated by Built In
Requisition Number: 2394338
Optum Insight is improving the flow of health data and information to create a more connected system. We remove friction and drive alignment between care providers and payers, and ultimately consumers. Our deep expertise in the industry and innovative technology empower us to help organizations reduce costs while improving risk management, quality and revenue growth. Ready to help us deliver results that improve lives? Join us in making healthcare work better for everyone through people-led, responsible AI while Caring. Connecting. Growing together.
The Optum Insight Engineering team is seeking a Lead Cloud Reliability Engineer with Site Reliability Engineering (SRE) experience to design, build, and scale modern AI Ops solutions across payer and provider transformation initiatives. This is a hands-on technical leadership role requiring deep involvement in architecture and engineering while leading globally distributed teams. The role focuses on delivering SRE Process automation, Infrastructure as code, AI, and Retrieval Augmented Generation (RAG) solutions with enterprise-grade reliability, security, and responsible AI practices.
You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities:
  • Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform, Python, Cloud Infrastructure while adhering to SRE best practices
  • Automate SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments through automation
  • Lead end-to-end design and implementation of Automated Ops solutions from concept through production with an emphasis on responsible AI practices
  • Design, code, test, and operate software using Python or Node.js
  • Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions
  • Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers
  • Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency

You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications:
  • 8+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure)
  • 3+ years of demonstrated hands-on automation experience with Python and Terraform based development
  • 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE)
  • Available to work rotating 24x7 primary and secondary on-call shifts

Preferred Qualifications:
  • Bachelor's degree in Computer Science, Software Engineering, or a related technical field (or 8+ years of equivalent software engineering experience in lieu of degree)
  • 2+ years delivering AI/ML or Generative AI solutions in production
  • CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform
  • Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
  • Experience leading globally distributed technical teams
  • Proven experience building Retrieval Augmented Generation (RAG) pipelines, Agentic AI, or multi-step AI workflows
  • Proven effective communication skills with ability to explain complex technical concepts to diverse stakeholders

*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $112,700 - $193,200 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.
At UnitedHealth Group, our mission is to help people live healthier lives and help make the health system work better for everyone. Together, we are shaping the future of healthcare by harnessing technology and innovation to make care simpler to navigate, more affordable and more connected for the people we serve. We are committed to creating an inclusive workplace where everyone feels welcomed, valued, heard and respected, empowering people to bring their authentic selves to work and strengthening our collective impact through diverse talents, backgrounds, experiences and perspectives.
UnitedHealth Group and its affiliated brands are an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.
UnitedHealth Group and its affiliated brands are a drug-free workplace. Candidates are required to pass a drug test before beginning employment.

Skills Required

  • 8+ years of overall software engineering and Site Reliability Engineering experience in a public cloud environment such as GCP, AWS, or Azure
  • 3+ years of hands-on automation experience with Python and Terraform-based development
  • 2+ years of experience managing Kubernetes environments, including EKS, AKS, or GKE
  • Availability to work rotating 24x7 primary and secondary on-call shifts
  • Bachelor's degree in Computer Science, Software Engineering, or a related technical field, or 8+ years of equivalent software engineering experience in lieu of a degree
  • 2+ years delivering AI, machine learning, or Generative AI solutions in production
  • CI/CD and Infrastructure as Code experience with Terraform
  • Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
  • Experience leading globally distributed technical teams
  • Experience building Retrieval Augmented Generation pipelines, Agentic AI, or multi-step AI workflows
  • Effective communication skills and ability to explain complex technical concepts to diverse stakeholders

What the Team is Saying

Optum Compensation & Benefits Highlights

  • Leave & Time Off Breadth — PTO is generally described as decent or good, and many note it as a strong part of the package. Actual ability to take time off can depend on workload and team coverage.
  • Retirement Support — Offerings include a 401(k) with employer match and access to an employee stock purchase plan, which are highlighted as meaningful components of total rewards. These programs are consistently referenced among core benefits.
  • Parental & Family Support — Parental and caregiver leave, along with adoption assistance, are publicly highlighted and viewed as notable elements of the package. Availability can be role-specific, but these supports contribute to overall breadth where offered.

Optum Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Eden Prairie, MN
160,000 Employees
Year Founded: 2011

What We Do

Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.

Gallery

Gallery
Gallery
Gallery

Optum Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.

Typical time on-site: Not Specified
HQEden Prairie, MN
Metro Manila, Philippines
Cebu, Philippines
Davao, Philippines
Ann Arbor, MI
Atlanta, GA
Baltimore, MD
Bengaluru, India
Chennai, India
Dallas, TX
Detroit, MI
Dublin, Ireland
Hartford, CT
Houston, TX
Hyderabad, India
Jacksonville, FL
Las Vegas, NV
Letterkenny, Ireland
Louisville, KY
Madison, WI
Minneapolis, MN
Nashville, TN
New Delhi, India
Philadelphia, PA
Phoenix, AZ
Pune, India
Raleigh, NC
San Diego, CA
Washington, DC
Learn more

Similar Jobs

Optum Logo Optum

Data Analyst

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
113K-193K Annually

Optum Logo Optum

Manager, Payer Advisory Services - Remote

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
113K-193K Annually

Optum Logo Optum

Senior Business Analyst

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
73K-130K Annually

Optum Logo Optum

Senior Manager, AI Innovation and Strategy - Life Sciences - Remote

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
113K-193K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account