Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.
We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and Operational Excellence initiatives across Optum bank. This leader will be responsible for improving service reliability, operational resiliency, deployment automation, observability, incident management, and production readiness for critical banking platforms.
The ideal candidate combines strong technical expertise with operational leadership experience, driving engineering excellence, automation, reliability, and continuous improvement while ensuring technology services meet business, customer, regulatory, and operational expectations.
This role will also help identify and implement emerging automation and AI-enabled operational capabilities that improve service health, reduce operational toil, and accelerate engineering productivity.
Primary Responsibilities:
- Leadership & Engineering Management
- Lead and develop multidisciplinary teams responsible for Site Reliability Engineering, DevOps, Platform Engineering, ITSM, and Operational Excellence
- Establish and execute enterprise reliability, availability, resiliency, and operational maturity strategies
- Drive engineering excellence through automation, observability, operational readiness, and continuous improvement practices
- Partner with Technology, Operations, Security, Infrastructure, Risk, and Business leaders to improve service reliability and customer experience
- Build and mentor high-performing teams while fostering accountability, innovation, operational ownership, and learning
- Manage staffing, capacity planning, talent development, succession planning, and organizational growth
- Establish operational metrics, governance standards, and service review processes to improve service performance and risk management
- SRE, DevOps & Platform Engineering
- Lead enterprise SRE practices including SLI/SLO adoption, error-budget management, reliability engineering, and operational maturity assessments
- Drive DevOps transformation initiatives, emphasizing automation, deployment standardization, CI/CD pipelines, Infrastructure-as-Code, and GitOps practices
- Establish production readiness standards and operational acceptance criteria for new technology deployments
- Improve platform resiliency through capacity planning, disaster recovery, fault tolerance, and resilience testing
- Drive reduction of operational toil through automation and self-healing capabilities
- Partner with application and infrastructure teams to improve system scalability, availability, and performance
- Lead initiatives to improve deployment frequency, reduce change failure rates, and accelerate service recovery times
- ITSM & Operational Excellence
- Establish and mature Incident, Problem, Change, Release, and Service Request Management processes
- Lead major incident management programs and executive communications during critical service disruptions
- Drive root-cause analysis and problem-management practices to eliminate recurring incidents
- Improve operational scorecards, service health reviews, and reliability reporting for executive stakeholders
- Ensure compliance with regulatory, audit, risk, and operational governance requirements
- Partner with Technology and Business leaders to improve service quality and customer outcomes through data-driven operational improvements
- Champion a culture of operational excellence and continuous service improvement
- Intelligent Automation & AI-Enabled Operations
- Identify opportunities to leverage AI and automation to improve operational effectiveness and engineering productivity
- Lead implementation and evaluation of solutions involving: AIOps, Intelligent alert correlation, Automated incident triage , Root cause analysis assistance, Knowledge management copilots, Agentic operational workflows etc.
- Partner with enterprise AI teams to evaluate emerging technologies that improve reliability and operational efficiency
- Drive responsible adoption of AI-enabled engineering and operational practices
- Support proof-of-concept initiatives that demonstrate measurable reductions in operational effort and incident resolution times
- Cross-Functional Leadership
- Collaborate with Engineering, Infrastructure, Security, Architecture, Risk, Compliance, and Operations teams to prioritize reliability and operational improvements
- Serve as a trusted advisor on reliability engineering, operational excellence, and automation strategies
- Drive alignment between technology and business stakeholders to improve service quality and operational outcomes
- Influence technology investment decisions that improve platform stability, resiliency, and operational efficiency
You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications:
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field
- 10+ years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or Technology Operations
- 5+ years of experience leading engineering or operational teams
- Proven experience supporting large-scale, business-critical production environments
- Experience with:
- SRE principles and practices
- DevOps and CI/CD
- ITSM processes
- Cloud platforms (Azure, AWS)
- On-prem environments
- Infrastructure-as-Code
- Container platforms (Kubernetes, OpenShift)
- Observability and monitoring platforms
- Experience with:
- Incident Management
- Problem Management
- Change Management
- Disaster Recovery
- Business Continuity
- Service Reliability Programs
- Experience leading operational transformations and continuous-improvement initiatives
Preferred Qualifications:
- Experience in banking, financial services, healthcare, or other highly regulated industries
- Experience implementing enterprise observability solutions such as Datadog, Splunk, Grafana, Prometheus, or OpenTelemetry
- Experience with cloud-native architectures and platform engineering practices
- Experience deploying AIOps, ChatOps, or intelligent automation solutions
- Familiarity with:
- Agentic AI workflows
- LLM-powered operational tooling
- Knowledge management platforms
- AI-enabled incident management solutions
- Experience establishing SLO frameworks and reliability governance programs
- Leadership Competencies
- Strategic thinker with solid operational and technology acumen
- Proven ability to build, lead, and inspire high-performing engineering and operational teams
- Solid understanding of service reliability, operational risk, resiliency, and governance
- Exceptional stakeholder management and executive communication skills
- Data-driven decision maker with a focus on measurable outcomes
- Solid execution and delivery leadership in complex enterprise environments
- Collaborative leader focused on continuous improvement, operational excellence, and customer outcomes
*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy.
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $112,700 to $193,200 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.
UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.
Skills Required
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field
- 10+ years in Software Engineering, Site Reliability, Platform Engineering, DevOps, Infrastructure, or Technology Operations
- 5+ years leading engineering or operational teams
- Proven experience supporting large-scale, business-critical production environments
- SRE principles and practices (SLI/SLO, error budget, reliability engineering)
- DevOps and CI/CD pipeline experience
- ITSM processes (Incident, Problem, Change, Release, Service Request Management)
- Cloud platform experience (Azure, AWS)
- On-premises environment experience
- Infrastructure-as-Code experience
- Container platforms (Kubernetes, OpenShift)
- Observability and monitoring platforms experience
- Incident Management, Problem Management, Change Management experience
- Disaster Recovery and Business Continuity experience
- Experience leading operational transformations and continuous-improvement initiatives
- Experience with service reliability programs and production readiness standards
- Proven stakeholder management and executive communication skills
- Experience in regulated industries (banking, financial services, healthcare)
- Experience with enterprise observability tools (Datadog, Splunk, Grafana, Prometheus, OpenTelemetry)
- Experience deploying AIOps, ChatOps, or intelligent automation solutions
- Familiarity with agentic AI workflows, LLM-powered tooling, and knowledge management platforms
- Experience establishing SLO frameworks and reliability governance programs
Optum Compensation & Benefits Highlights
-
Healthcare Strength — Official materials highlight copay and HSA medical plan choices with in‑network preventive care at 100%, prescription coverage, and low/no‑cost virtual visits, plus company HSA contributions. Dental preventive services are 100% in network, and mental health resources include an EAP and premium Calm access.
-
Parental & Family Support — Programs include six weeks paid parental leave, up to two weeks paid caregiver leave, and Bright Horizons back‑up care with enhanced family supports. Adoption assistance up to $10,000 for full‑time employees reinforces family‑oriented benefits.
-
Equity Value & Accessibility — Financial benefits include an Employee Stock Purchase Plan at a 10% discount, expanding access to equity ownership.
Optum Insights
What We Do
Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.
Gallery
Optum Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.
Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.