DevOps & Infrastructure Engineer (HPC/GPU)

Reposted 13 Days Ago
2 Locations
Remote or Hybrid
Entry level
Information Technology • Software
The Role
Own and develop HPC/GPU infrastructure for R&D and customer deployments. Design cross-platform CI/CD pipelines, manage on-premise and cloud Kubernetes clusters with GPU passthrough and MIG, configure self-hosted GitHub Actions runners, and maintain DockerHub registries. Assemble and tune physical GPU servers, manage PCIe topology and hardware constraints, and maintain compatibility across NVIDIA drivers, CUDA, and ROCm versions. Provide customer-facing technical assistance for GPU container deployments.
Summary Generated by Built In
Your Mission

You will own the infrastructure that powers our R&D and helps our customers deploy our technology on-premise. You will move beyond standard cloud DevOps into the world of High-Performance Computing (HPC).

  • Think: Design a robust CI/CD strategy that handles cross-platform compilation (Windows/Linux) and execution on specific hardware targets (NVIDIA A100, AMD MI250, Consumer GPUs). Architect solution templates for our customers who need to deploy Hybridizer-generated binaries on their own private clouds.

  • Implement:

    • Set up and maintain Kubernetes clusters (both on-premise and cloud) with GPU Passthrough and Multi-Instance GPU (MIG) configurations.

    • Develop GitHub Actions pipelines that seamlessly dispatch heavy test suites to self-hosted runners equipped with specific GPU accelerators.

    • Configure DockerHub registries and secure container lifecycles for our compiler images.

  • Build:

    • Hardware Tuning: Assemble and fine-tune physical servers. This includes managing PCIe topology, cooling profiles, and power constraints to ensure consistent benchmarking results.

    • Driver Ecosystem: Manage the complex matrix of NVIDIA drivers, CUDA toolkits, and ROCm versions across our fleet, ensuring compatibility with our compiler’s output.

What You Bring to the Table

You are a DevOps engineer who loves hardware. You understand that "the cloud" is just someone else's computer, and sometimes you need to manage that computer yourself.

  • Core DevOps: Strong mastery of Docker and Kubernetes. You know how to write custom Helm charts and manage stateful sets.

  • GPU Infrastructure: You have hands-on experience with NVIDIA Container Toolkit or ROCm integration in containers. You understand concepts like PCIe passthrough, IOMMU groups, and GPU orchestration.

  • CI/CD Automation: Expert in GitHub Actions. You can write complex workflows with matrix strategies and self-hosted runners.

  • System Administration: You are comfortable with Linux kernel tuning, driver installation (dkms), and diagnosing hardware bottlenecks.

  • Customer Facing: You have the communication skills to assist clients. You can explain how to expose a GPU to a Docker container to a sysadmin who might not be an expert in HPC.

  • Adaptability: You are ready to work with a mix of consumer and data-center grade hardware (e.g., configuring a server with 4x RTX 5090s or managing a DGX station).

Skills Required

  • Strong mastery of Docker and Kubernetes
  • Experience writing custom Helm charts and managing StatefulSets
  • Hands-on experience with NVIDIA Container Toolkit or ROCm integration in containers
  • Understanding of PCIe passthrough, IOMMU groups, and GPU orchestration
  • Expertise with GitHub Actions, including matrix strategies and self-hosted runners
  • Comfort with Linux kernel tuning, driver installation using DKMS, and hardware troubleshooting
  • Communication skills for assisting customers with GPU container deployments
  • Ability to work with consumer and data-center GPU hardware
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1 Employee
Year Founded: 2008

What We Do

Hybridizer is a software platform for performance portability and GPU acceleration. Its compiler transforms C#/.NET and Java bytecode or high-level code into optimized source code for multicore CPUs and GPUs, allowing developers to use existing codebases without learning CUDA or rewriting applications. The technology supports debugging, profiling, cross-platform deployment, and demanding workloads such as quantitative finance, scientific simulation, and data processing.

Similar Jobs

Mastercard Logo Mastercard

Director, Global Communications

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Harrison, NY, USA
38800 Employees
174K-286K Annually

Capital One Logo Capital One

Director, AI Engineering (Remote - eligible)

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
3 Locations
55000 Employees
245K-335K Annually

Capital One Logo Capital One

Director, Data Science - Consumer and Developer Experience (Remote-Eligible)

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
6 Locations
55000 Employees
245K-335K Annually

Optum Logo Optum

Senior Medical Director, Population Health - Remote

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Minneapolis, MN, USA
160000 Employees
292K-439K Annually

Similar Companies Hiring

Ford Energy Thumbnail
Automotive • Software • Energy • Utilities • Manufacturing • Renewable Energy
US
55 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account