Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.
We are seeking a Senior Platform DevOps Engineer with Site Reliability Engineering (SRE) experience to help build, automate, secure, and operate enterprise cloud platforms and services. This role will be responsible for designing and maintaining cloud infrastructure, deployment automation, observability solutions, security controls, operational processes, and platform standards that enable development teams to deliver secure, reliable, and scalable applications.
The ideal candidate is a hands-on engineer who thrives in complex cloud environments and has experience supporting production platforms, modern CI/CD practices, infrastructure as code, monitoring and alerting, cloud security, incident response, and operational excellence initiatives.
You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities:
- Design, build, and maintain cloud infrastructure using Infrastructure as Code (IaC) practices
- Develop, enhance, and support CI/CD pipelines and deployment automation across multiple environments
- Implement platform standards, reusable automation frameworks, and engineering best practices
- Support cloud-native application deployments and containerized workloads
- Implement and maintain monitoring, logging, alerting, and observability solutions to improve system reliability and operational visibility
- Define and support SRE practices including service health monitoring, operational readiness, incident response, runbooks, and reliability improvements
- Troubleshoot and resolve complex application, infrastructure, networking, identity, security, and deployment issues
- Manage and improve platform security by implementing security controls, vulnerability remediation processes, secret management, and automated security scans
- Partner with development, security, architecture, and operations teams to improve platform reliability, scalability, and delivery speed
- Support disaster recovery, resiliency, backup, recovery, and business continuity initiatives
- Participate in on-call and incident response activities as required
- Automate repetitive operational tasks using scripting and cloud automation tools
- Drive continuous improvement across DevOps, SRE, platform engineering, cloud operations, and governance practices
- Mentor engineers and promote DevOps and SRE best practices across teams
- Evaluate and leverage AI-assisted tools and automation to improve engineering productivity, operational efficiency, and software delivery processes
You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications:
- 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, Cloud Engineering, or related technology roles
- 4+ years of hands-on experience supporting production workloads in Microsoft Azure
- 4+ years of experience implementing Infrastructure as Code using Terraform
- 4+ years of experience designing and supporting CI/CD pipelines using GitHub Actions
- 1+ years of experience with container technologies such as Docker and Kubernetes
- 1+ years of experience with cloud networking, identity management, secrets management, and platform security
- 1+ years of experience implementing monitoring, logging, alerting, dashboards, and operational observability solutions
- 1+ years of experience supporting production incidents, troubleshooting complex technical issues, and driving root cause analysis
- 1+ years of experience remediating security vulnerabilities and implementing secure software delivery practices
- 1+ years of experience with Git-based source control, branching strategies, repository governance, and release management
- 1+ years of experience with scripting and automation using PowerShell, Python, Bash, or similar languages
Preferred Qualifications:
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field, or equivalent practical experience.
- Experience with Kubernetes platform administration and GitOps practices
- Experience with cloud-native monitoring and observability platforms
- Experience supporting highly available, mission-critical production systems
- Experience with application security scanning, dependency management, and software supply chain security
- Experience implementing SLOs, SLIs, error budgets, and other reliability engineering practices
- Experience with disaster recovery planning, resiliency testing, and capacity management
- Experience working in regulated enterprise environments
- Experience with AI-assisted development, platform automation, or operational tooling
- Prior experience serving as a technical lead, mentor, or senior engineering resource within a platform engineering, DevOps, or SRE organization
- Proven understanding of cloud infrastructure, reliability engineering, security, and operational excellence principles
*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy.
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $91,700 - $163,700 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.
UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.
Skills Required
- 5+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Engineering, or related technology roles
- 4+ years of hands-on experience supporting production workloads in Microsoft Azure
- 4+ years of experience implementing Infrastructure as Code using Terraform
- 4+ years of experience designing and supporting CI/CD pipelines using GitHub Actions
- 1+ years of experience with Docker and Kubernetes
- 1+ years of experience with cloud networking, identity management, secrets management, and platform security
- 1+ years of experience implementing monitoring, logging, alerting, dashboards, and operational observability solutions
- 1+ years of experience supporting production incidents, troubleshooting complex technical issues, and driving root cause analysis
- 1+ years of experience remediating security vulnerabilities and implementing secure software delivery practices
- 1+ years of experience with Git-based source control, branching strategies, repository governance, and release management
- 1+ years of experience with scripting and automation using PowerShell, Python, Bash, or similar languages
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field, or equivalent practical experience
- Experience with Kubernetes platform administration and GitOps practices
- Experience with cloud-native monitoring and observability platforms
- Experience supporting highly available, mission-critical production systems
- Experience with application security scanning, dependency management, and software supply chain security
- Experience implementing SLOs, SLIs, error budgets, and other reliability engineering practices
- Experience with disaster recovery planning, resiliency testing, and capacity management
- Experience working in regulated enterprise environments
- Experience with AI-assisted development, platform automation, or operational tooling
- Prior experience serving as a technical lead, mentor, or senior engineering resource within a platform engineering, DevOps, or SRE organization
- Proven understanding of cloud infrastructure, reliability engineering, security, and operational excellence principles
Optum Compensation & Benefits Highlights
-
Parental & Family Support — Paid parental leave (six weeks), paid caregiver leave (up to two weeks), Bright Horizons back-up care, and adoption assistance up to $10,000 are prominently included. Feedback suggests these family supports meaningfully aid work-life balance and are often highlighted as strengths.
-
Retirement Support — A 401(k) with company match is available to all employees, including part-time staff, alongside other financial protections like disability and life insurance. Feedback suggests broad access and matching make retirement support a core pillar of the package.
-
Equity Value & Accessibility — An Employee Stock Purchase Plan offers discounted company stock, with some roles also eligible for sign-on or performance bonuses. Feedback suggests the ESPP is a standout financial perk that helps employees build ownership over time.
Optum Insights
What We Do
Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.
Gallery
Optum Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.
Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.