Platform Engineer, AI/ML Infrastructure

Posted Yesterday
Hiring Remotely in Austin, Texas, USA
In-Office or Remote
120K-250K Annually
Mid level
Cloud • Software
The Role
Build and operate Kubernetes-based AI/ML infrastructure, including GPU scheduling, multi-tenant isolation, platform services, infrastructure as code, GitOps, CI/CD, observability, and reliability practices. Deploy reproducible platforms in cloud, restricted, and disconnected environments. Develop runbooks, support incident response, contribute to open-source infrastructure projects, and collaborate with security engineers and government stakeholders. The role requires production infrastructure experience, cloud expertise, programming skills, and eligibility for a U.S. security clearance.
Summary Generated by Built In
Who We Are

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.
OpenTeams exists to make ownership possible.
Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.
If that sounds like your kind of work, we'd like to meet you.

Platform Engineer, AI/ML Infrastructure

Location:  U.S - Remote OR Hybrid - Washington, DC, Denver, CO or Colorado Springs, CO. 

Work Authorization: U.S. citizenship required

Clearance:  U.S.-Remote Opening: An active clearance is not required. Candidates must be eligible and willing to obtain and maintain a U.S. security clearance. Hybrid Opening: An active TS/SCI clearance is preferred. Candidates may also be considered if they previously held a TS/SCI with CI polygraph or currently hold an active TS or Secret clearance.

Salary Range: $120,000–$250,000 USD, dependent on experience level and location

Openings: Two positions are available:

  • One hybrid position: Candidates must be located in or willing to work hybrid from Washington, DC; Denver, CO; or Colorado Springs, CO. An active TS/SCI clearance is required. Up to 15% travel is required.
  • One U.S.-remote position: Candidates may work remotely from anywhere in the United States. An active clearance is not required, but candidates must be willing and able to undergo the process required to obtain and maintain a U.S. security clearance.

Candidates will be considered for the opening that best aligns with their location, clearance status, experience, and work preferences. Candidates who meet the requirements for multiple openings may be considered for more than one.

Candidates will be considered for the opening that best aligns with their location, clearance status, experience, and work preferences. Candidates who meet the requirements for multiple openings may be considered for more than one.

About the Role

OpenTeams builds AI platforms that governments and enterprises own outright: the infrastructure, the data, the models, and the evidence that the whole thing does what it claims. We're hiring several engineers to build and run that infrastructure.

The work spans the full depth of an AI platform. The Kubernetes clusters that schedule GPU workloads, move large datasets, and keep tenants isolated from one another. The services that make it a platform rather than a cluster: workflow orchestration, data ingest, model serving, policy enforcement, audit logging. The delivery path that gets released into production reliably and can prove what it shipped. The cloud infrastructure that ties it all together, and the operational practices that keep a distributed system resilient.

Much of this has to run where you can't assume normal cloud resources, or even an internet connection. That constraint is the interesting part of the job. Portability, reproducibility, and operability are design inputs from the first commit rather than problems handed downstream.

We build on open source and contribute back. Kubernetes, Terraform and OpenTofu, Argo, Prometheus, Nebari, and others. Upstream work is part of the job, not something you do on weekends.

This posting covers multiple roles, spanning mid-level through senior. We understand nobody spans every area above, so tell us where you fit. We assign level-based roles based on what you've actually done rather than a year count.

Key Responsibilities
  • Build and operate Kubernetes-based infrastructure for demanding AI/ML workloads, including GPU scheduling, resource management, and multi-tenant isolation
  • Design and implement platform services for orchestration, data ingest, model serving, and results management behind documented APIs
  • Write infrastructure as code and build GitOps pipelines so environments are reproducible from source
  • Build and operate CI/CD pipelines that produce versioned, signed, scanned release artifacts along with the documentation needed to deploy them
  • Own reliability: capacity planning, upgrade paths, failure-mode analysis, backup and recovery, incident response, and postmortems
  • Implement monitoring, logging, tracing, and alerting, and define the service level objectives they're measured against
  • Deploy and validate the platform in restricted, disconnected, or limited-connectivity environments, and verify parity after each release
  • Keep the platform portable by constraining dependencies to what's confirmed available in target environments
  • Write runbooks and operational documentation that other engineers can execute without you in the room
  • Contribute to Nebari and other open-source infrastructure, Kubernetes, and MLOps projects
  • Work with security engineers, government stakeholders, and other engineers to turn requirements into systems that hold up
  • Collaborate asynchronously across a distributed team
Required Skills & Experience
  • U.S. citizenship, and the ability to obtain and maintain a U.S. security clearance
  • Four or more years of hands-on experience building or operating production infrastructure, platforms, or distributed systems
  • Production experience with Kubernetes and containerized workloads
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud
  • Experience with infrastructure as code and CI/CD, using tools such as Terraform, OpenTofu, Pulumi, or Helm
  • Working proficiency in Python, Go, Bash, or a comparable language
  • Experience implementing or operating production monitoring and observability
  • Ability to write documentation, runbooks, and deployment procedures that other people can actually follow
  • Ability to work independently and collaborate well in a remote, distributed team
Nice to Have

You will not have all of these, and very few people will. They're the things that would help, not a checklist. If the required list above describes you, apply.

  • An active U.S. security clearance, particularly TS/SCI with CI polygraph
  • Experience deploying or operating software in air-gapped, disconnected, or otherwise restricted environments
  • Experience with Department of Defense, Intelligence Community, or comparably regulated programsFamiliarity with the Risk Management Framework, NIST 800-53 or 800-171, or similar frameworks, and with producing the evidence they require
  • Experience supporting an Authorization to Operate, or with continuous ATO modelsSupply chain security work: hardened images, artifact signing, SBOM generation, dependency and container scanning, policy enforcement
  • Familiarity with cross-domain solutions, guards, data diodes, or similar transfer mechanisms
  • Experience with classified cloud environments, including AWS Secret or Top Secret regionsA DoD 8140/8570 qualifying certification such as Security+, CISSP, CASP+, or CISM, or willingness to obtain one after joining
  • Experience building MLOps pipelines or infrastructure for AI/ML workloads
  • Experience with GPU scheduling, distributed inference, or large-scale data and evaluation pipelines
  • Experience with model-serving or gateway frameworks such as KServe, vLLM, or LLM-DExperience designing API-first services and vendor-agnostic platforms that run across multiple environments
  • Experience with agentic workflow frameworks or multi-step AI pipeline orchestration
  • Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects, and experience with Nebari specifically
  • Familiarity with data sovereignty and privacy requirements for enterprise or government AI systems
  • Experience leading technical initiatives, setting engineering standards, or mentoring other engineers
  • Experience supporting rapid prototyping programs or defense innovation initiatives

Grow With Us

At OpenTeams, growth isn’t just about the company—it’s about you.
We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.

Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily.  That global perspective and diversity makes our solution more universal and robust.  We are committed to continuing to celebrate diversity on our team.

Supported people are successful people.  We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work.
We invest  in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration.

Commitment to diversity, equity, inclusion, and belonging

OpenTeams understands that valuing diverse creative practices and forms of knowledge is crucial to and enriches the company’s core mission. We encourage applications from everyone, including members of all equity-seeking communities, such as (but certainly not limited to) women, racialized and Indigenous persons, disabled people, persons of all sexual orientations, gender identities and expressions.

We are an equal opportunity employer - all qualified applicants will receive equal consideration for recruitment, interviews, employment, training, compensation, promotion, and related activities. We do not discriminate based on race, religion, gender, gender identity, gender expression, color, national origin, pregnancy, ancestry, domestic partner status, disability, sexual orientation, age, genetic predisposition, medical condition, marital status, citizenship status, military or veteran status, or any other basis covered by applicable laws. OpenTeams will not tolerate discrimination or harassment based on these characteristics or any other unlawful behavior, conduct, or purpose.


Skills Required

  • U.S. citizenship
  • Ability to obtain and maintain a U.S. security clearance
  • Four or more years of hands-on experience building or operating production infrastructure, platforms, or distributed systems
  • Production experience with Kubernetes and containerized workloads
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud
  • Experience with infrastructure as code and CI/CD using Terraform, OpenTofu, Pulumi, or Helm
  • Working proficiency in Python, Go, Bash, or a comparable language
  • Experience implementing or operating production monitoring and observability
  • Ability to write documentation, runbooks, and deployment procedures
  • Ability to work independently and collaborate effectively in a remote, distributed team
  • Active U.S. security clearance, particularly TS/SCI with CI polygraph
  • Experience with air-gapped, disconnected, or restricted environments
  • Experience supporting Department of Defense, Intelligence Community, or regulated programs
  • Familiarity with RMF, NIST 800-53, NIST 800-171, or similar frameworks
  • Experience supporting an Authorization to Operate or continuous ATO models
  • Supply chain security experience, including hardened images, artifact signing, SBOMs, scanning, and policy enforcement
  • Familiarity with cross-domain solutions, guards, or data diodes
  • Experience with classified cloud environments
  • DoD 8140/8570 certification such as Security+, CISSP, CASP+, or CISM, or willingness to obtain one
  • Experience building MLOps pipelines or AI/ML infrastructure
  • Experience with GPU scheduling, distributed inference, or large-scale data and evaluation pipelines
  • Experience with model-serving or gateway frameworks such as KServe, vLLM, or LLM-D
  • Experience designing API-first, vendor-agnostic platforms across multiple environments
  • Experience with agentic workflow frameworks or multi-step AI pipeline orchestration
  • Open-source contributions to Kubernetes, infrastructure, MLOps, or observability projects, particularly Nebari
  • Familiarity with data sovereignty and privacy requirements for enterprise or government AI systems
  • Experience leading technical initiatives, setting engineering standards, or mentoring engineers
  • Experience supporting rapid prototyping or defense innovation initiatives
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Austin, Texas
36 Employees
Year Founded: 2019

What We Do

OpenTeams is at the forefront of open source support, offering a wide range of practice areas led by a network of Open Source Architects. With over 680 open source technologies, our team provides comprehensive services including strategy and consulting, custom development, integration, migration, and 24/7 support. Our practice areas cover various domains, such as Machine Learning Operations, Cloud Optimization, Data Science and Engineering, SaaS and Cloud Applications, Artificial Intelligence and Machine Learning, PyTorch Hardware Optimization, PyTorch Artificial Intelligence System Building, and High-Performance Systems. Each solution is staffed by experienced professionals who assist businesses in addressing specific challenges and leveraging open source technologies to achieve their goals. OpenTeams is dedicated to helping clients build better software with reliable open source support.

Similar Jobs

Deepgram Logo Deepgram

Platform Engineer

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Remote
USA
150 Employees
136K-240K Annually

Rapid7 Logo Rapid7

Senior Director, Customer Innovation

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
211K-285K Annually

Rapid7 Logo Rapid7

Vector Command Specialist

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
89K-121K Annually

Wipfli Logo Wipfli

Consultant

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
2900 Employees
117K-158K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account