Senior Software Engineer, AI Infrastructure

Reposted Yesterday
Hiring Remotely in US
Remote
Senior level
Software
The Role
Design and build Kubernetes-based LLM serving infrastructure, including GPU scheduling, deployment, scaling, model lifecycle management, Helm packaging, offline enterprise installs, API integrations, identity, metering, and production observability. Contribute across a multi-service codebase, develop GPU telemetry, write technical design documents, and guide engineering direction on a remote-first senior team.
Summary Generated by Built In
Company Description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

Mirantis is building a new enterprise AI infrastructure product that lets organizations run and govern large language models on their own Kubernetes clusters. You will join a small senior team early, with broad ownership of the model-serving layer and its path to production.

What you'll do

  • Design and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management.

  • Package the platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks.

  • Integrate the serving layer with the platform's API gateway, identity, and metering services.

  • Build the observability for operating GPU inference in production (serving metrics, GPU telemetry).

  • Contribute across a multi-service codebase and help set engineering direction through design docs and reviews.

Qualifications

What we're looking for:

  • 5+ years of software engineering experience in infrastructure, platform, or distributed systems.

  • Deep hands-on Kubernetes experience: building and operating production workloads and Helm charts, not just consuming managed clusters.

  • Experience with GPU workloads or LLM inference, or strong adjacent systems experience and a track record of learning fast.

  • Strong Go programming skills; solid CI/CD and infrastructure-as-code skills.

  • Fluency with AI-assisted development tools (Claude Code, OpenAI Codex) as part of your daily engineering workflow.

  • Comfortable with high autonomy on a small, remote-first, written-culture team.

Nice to have

  • Inference performance work (quantization, batching, caching) or distributed serving frameworks.

  • Enterprise deployment experience: air-gapped installs, SSO/OIDC, supply-chain security.

  • UI development experience (e.g. React/TypeScript), useful as the product's management surfaces grow.

  • Open-source contributions in the Kubernetes or ML-infrastructure ecosystems.

 

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Skills Required

  • 5+ years of software engineering experience in infrastructure, platform, or distributed systems
  • Deep hands-on Kubernetes experience building and operating production workloads
  • Experience creating and operating Helm charts
  • Experience with GPU workloads or LLM inference, or strong adjacent systems experience with demonstrated ability to learn quickly
  • Strong Go programming skills
  • Solid CI/CD and infrastructure-as-code skills
  • Fluency with AI-assisted development tools such as Claude Code or OpenAI Codex
  • Ability to work autonomously on a small, remote-first, written-culture team
  • Inference performance experience with quantization, batching, caching, or distributed serving frameworks
  • Enterprise deployment experience with air-gapped installations, SSO/OIDC, or supply-chain security
  • UI development experience with React or TypeScript
  • Open-source contributions in Kubernetes or ML infrastructure ecosystems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Campbell, CA
729 Employees
Year Founded: 1999

What We Do

We are dedicated to helping organizations increase developer productivity and ship code faster on public and private clouds. We provide a ZeroOps experience to remove the stress of managing cloud native infrastructure by combining software and automation tools with our cloud native expertise to deliver the industry's leading secure cloud platforms. Our capabilities allow us to provide a secure and reliable cloud native platform that includes validated FIPS-140-2 Encryption and DISA STIG ready capabilities. Who do we serve? We serve a wide range of industries, building on our extensive customer experience to provide distinct value in specific verticals including Financial Services, Government & Education, Healthcare, Manufacturing, and Telecommunications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Inmarsat, PayPal, Reliance Jio, Societe Generale, Splunk, and S&P Global. Learn more at www.mirantis.com.

Similar Jobs

NVIDIA Logo NVIDIA

Senior Technical Marketing Engineer - DSX AI Infrastructure Software

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Remote or Hybrid
2 Locations
21960 Employees
160K-322K Annually

WorkHero Logo WorkHero

Senior Software Engineer

Artificial Intelligence • Software • Automation
Remote
USA
27 Employees

NVIDIA Logo NVIDIA

Software Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
3 Locations
21960 Employees
184K-357K Annually

Coinbase Logo Coinbase

Senior Software Engineer

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
USA
4700 Employees
186K-219K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account