Principal AI/ML Platform Engineer - Remote

Posted 3 Hours Ago
Hiring Remotely in Eden Prairie, MN, USA
In-Office or Remote
165K-282K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
The Role
Own the architecture and long-term evolution of hybrid, multi-tenant AI compute platforms across bare-metal OpenShift and public-cloud services. Define distributed training networking, GPU utilization and cost models, GitOps governance, model-serving standards, workload placement, identity and security controls, SLOs, disaster recovery, and platform upgrade strategies. Partner with AI, privacy, and security teams to deliver HIPAA-compliant infrastructure for training and inference workloads.
Summary Generated by Built In
Requisition Number: 2386255
Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.
As a Principal AI/ML Platform Engineer on the UnitedHealth Group (UHG) enterprise team, you will serve as the AI Architect across IaaS and PaaS environments, owning the technical direction, reference architecture, and long-term evolution of our multi-tenant AI compute platform end to end. Our team builds and maintains an advanced compute estate spanning on-premises bare-metal Red Hat OpenShift AI clusters equipped with high-performance NVIDIA GPUs and InfiniBand/RoCE training fabrics, alongside public-cloud managed AI services including Azure AI Foundry, AWS Bedrock, and GCP Vertex AI. In this role, you will define architecture standards, optimize high-throughput model training and inference pipelines, establish cost and utilization economics, and enforce strict HIPAA, security, and data-governance standards for regulated healthcare workloads.
You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities:
  • Own the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid on-premises bare-metal OpenShift AI clusters and public-cloud managed AI platforms (Azure AI Foundry, AWS Bedrock, GCP Vertex AI)
  • Set network, latency, and topology standards for distributed training (NVLink, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL), ensuring interconnect boundaries are strictly maintained
  • Establish cluster governance, GitOps workflows (Argo CD), RHACM policies, and automated lifecycle management for bare-metal accelerated compute nodes
  • Design cost and utilization models including capex amortization, accelerator-sharing strategies (MIG/time-slicing), and cost-per-token/training-run math to inform accelerator procurement roadmaps
  • Standardize model-serving platforms (vLLM, KServe) and inference gateways, setting quantization policies, provenance review gates, and intelligent model routing rules
  • Define workload placement frameworks to determine self-hosted versus managed cloud deployment based on data residency, latency, cost, and compliance requirements
  • Design and enforce identity, access, and security controls for AI workloads and autonomous agents, including least-privilege RBAC, Vault secret management, short-lived credentials, and mTLS
  • Establish enterprise Service Level Objectives (SLOs), disaster recovery plans, and upgrade strategies for OpenShift, OpenShift AI, GPU operators, drivers, and firmware
  • Partner with cross-functional AI teams, LLM gateway engineers, privacy, and security stakeholders to ensure seamless integration and HIPAA compliance

You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications:
  • Bachelor's degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree
  • 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering
  • 5+ years of experience operating Kubernetes or OpenShift at scale in production bare-metal or enterprise cloud environments
  • 3+ years of experience designing and managing accelerated-compute (AI/GPU) infrastructure utilizing NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM
  • 3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services (e.g., Azure AI Foundry, AWS Bedrock, or GCP Vertex AI)
  • 3+ years of experience with HPC/AI networking technologies, including InfiniBand or RoCEv2, GPUDirect RDMA, and NCCL collective communication limits
  • 3+ years of experience managing multi-cluster fleets using RHACM (or equivalent) and GitOps tooling (Argo CD or Flux)
  • 3+ years of experience implementing enterprise security and IAM controls for software workloads (RBAC, OIDC/OAuth, Vault secrets management, mTLS)

Preferred Qualifications:
  • Experience with distributed training frameworks (PyTorch DDP/FSDP, DeepSpeed, Ray, JAX) and batch scheduling systems (Kueue, Volcano)
  • Hands-on experience with LLM inference serving technologies (vLLM, TensorRT-LLM, KServe) and platform tooling such as OpenShift AI (RHOAI) or Kubeflow pipelines
  • Experience with high-performance parallel storage systems (Ceph/ODF, Lustre, IBM Storage Scale, VAST, WEKA)
  • Active Red Hat certifications (e.g., Red Hat Certified Architect / RHCA) or open-source contributions to CNCF, OpenShift, or AI infrastructure projects
  • Experience operating AI/ML platforms within regulated healthcare environments under HIPAA and UHG data privacy controls

*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy.
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $164,600 - $282,200 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.
UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.
#optumtechpj

Skills Required

  • Bachelor's degree or 4+ years of equivalent software or platform engineering experience in lieu of a degree
  • 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering
  • 5+ years of experience operating Kubernetes or OpenShift at scale in production bare-metal or enterprise cloud environments
  • 3+ years of experience designing and managing accelerated-compute AI/GPU infrastructure using NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM
  • 3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services
  • 3+ years of experience with HPC/AI networking technologies including InfiniBand or RoCEv2, GPUDirect RDMA, and NCCL
  • 3+ years of experience managing multi-cluster fleets using RHACM or equivalent and GitOps tooling such as Argo CD or Flux
  • 3+ years of experience implementing enterprise security and IAM controls including RBAC, OIDC/OAuth, Vault secrets management, and mTLS
  • Experience with distributed training frameworks such as PyTorch DDP/FSDP, DeepSpeed, Ray, or JAX and batch scheduling systems such as Kueue or Volcano
  • Hands-on experience with LLM inference serving technologies such as vLLM, TensorRT-LLM, or KServe and platform tooling such as OpenShift AI or Kubeflow Pipelines
  • Experience with high-performance parallel storage systems such as Ceph/ODF, Lustre, IBM Storage Scale, VAST, or WEKA
  • Active Red Hat certifications such as Red Hat Certified Architect/RHCA or open-source contributions to CNCF, OpenShift, or AI infrastructure projects
  • Experience operating AI/ML platforms in regulated healthcare environments under HIPAA and UHG data privacy controls

What the Team is Saying

Optum Compensation & Benefits Highlights

  • Parental & Family Support Paid parental leave (six weeks), paid caregiver leave (up to two weeks), Bright Horizons back-up care, and adoption assistance up to $10,000 are prominently included. Feedback suggests these family supports meaningfully aid work-life balance and are often highlighted as strengths.
  • Retirement Support A 401(k) with company match is available to all employees, including part-time staff, alongside other financial protections like disability and life insurance. Feedback suggests broad access and matching make retirement support a core pillar of the package.
  • Equity Value & Accessibility An Employee Stock Purchase Plan offers discounted company stock, with some roles also eligible for sign-on or performance bonuses. Feedback suggests the ESPP is a standout financial perk that helps employees build ownership over time.

Optum Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Eden Prairie, MN
160,000 Employees
Year Founded: 2011

What We Do

Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.

Gallery

Gallery
Gallery
Gallery

Optum Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.

Typical time on-site: Not Specified
HQEden Prairie, MN
Metro Manila, Philippines
Cebu, Philippines
Davao, Philippines
Ann Arbor, MI
Atlanta, GA
Baltimore, MD
Bengaluru, India
Chennai, India
Dallas, TX
Detroit, MI
Dublin, Ireland
Hartford, CT
Houston, TX
Hyderabad, India
Jacksonville, FL
Las Vegas, NV
Letterkenny, Ireland
Louisville, KY
Madison, WI
Minneapolis, MN
Nashville, TN
New Delhi, India
Philadelphia, PA
Phoenix, AZ
Pune, India
Raleigh, NC
San Diego, CA
Washington, DC
Learn more

Similar Jobs

Optum Logo Optum

Consultant

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
92K-164K Annually

Optum Logo Optum

Senior Product Manager

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
113K-193K Annually

Optum Logo Optum

Consultant

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
113K-193K Annually

Optum Logo Optum

Lead Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
113K-193K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account