Cluster Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in USA
Remote
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
The Role
Architect, deploy, optimize, and operate large-scale GPU clusters for AI training and inference. Responsibilities include distributed PyTorch and NCCL tuning, GPU networking and storage optimization, scheduling, benchmarking, troubleshooting, automation, monitoring, and performance engineering across compute, networking, storage, and software layers. The role supports production environments with hundreds to thousands of GPUs and partners with ML engineers to improve training scalability and inference efficiency.
Summary Generated by Built In

Position Summary
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization.
The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.
Responsibilities

  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.

  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.

  • Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.

  • Build and support production AI infrastructure running hundreds to thousands of GPUs.

  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.

  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.

  • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.

  • Configure and tune distributed AI software stacks including:

    • PyTorch

    • NCCL

    • CUDA

    • UCX

    • MPI

    • Slurm

    • Pyxis/Enroot

  • Optimize GPU scheduling and resource allocation for both training and inference environments.

  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.

  • Identify performance regressions and troubleshoot distributed training issues at scale.

  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.

  • Work closely with ML engineers to improve training scalability and inference efficiency.

  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.

  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.


    Required Qualifications

  • 7+ years designing or operating large-scale Linux infrastructure.

  • 5+ years supporting production GPU clusters for AI or HPC workloads.

  • Demonstrated experience building multi-node GPU training environments from the ground up.

  • Deep expertise with distributed PyTorch training.

  • Extensive experience troubleshooting and optimizing NCCL communications.

  • Strong understanding of distributed AI communication patterns, including:

    • AllReduce

    • ReduceScatter

    • AllGather

    • Broadcast

    • Point-to-point communications

  • Experience benchmarking distributed training using tools such as:

    • nccl-tests

    • NVIDIA DCGM

    • Nsight Systems

    • MLPerf (preferred)

  • Strong understanding of GPU memory management, including:

    • KV Cache

    • Activation checkpointing

    • Tensor Parallelism

    • Pipeline Parallelism

    • Data Parallelism

  • Experience optimizing LLM inference throughput, including:

    • Tokens/sec optimization

    • Batch sizing

    • Continuous batching

    • KV cache tuning

    • Memory bandwidth optimization

  • Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.

  • Expert-level Linux systems administration skills.

  • Experience with Slurm workload manager.

  • Experience using Pyxis and Enroot for containerized GPU workloads.

  • Strong scripting skills using Python and Bash.


Technical Expertise
AI Frameworks

  • PyTorch

  • CUDA

  • NCCL

  • Triton (preferred)

  • TensorRT-LLM (preferred)

Cluster Scheduling

  • Slurm

  • Pyxis

  • Enroot

GPU Networking
Strong understanding of:

  • InfiniBand

  • RoCE v2

  • RDMA

  • GPUDirect RDMA

  • GPUDirect Storage

  • UCX

  • MPI

  • Network topology optimization

  • Congestion control

  • QoS

  • ECN/PFC

  • High-speed Ethernet (200/400/800 GbE)

Storage
Experience designing or tuning storage for AI workloads, including:

  • Parallel file systems

  • Distributed storage

  • Object storage

  • NVMe

  • Checkpoint optimization

  • Dataset staging

  • GPUDirect Storage

  • Storage bandwidth optimization

  • Metadata performance

Performance Engineering
Experience with:

  • NCCL benchmarking

  • Multi-node scaling analysis

  • GPU utilization optimization

  • Communication/computation overlap

  • NUMA optimization

  • CPU affinity

  • PCIe topology

  • GPU topology (NVLink/NVSwitch)

  • Memory bandwidth analysis

  • End-to-end performance profiling


Preferred Qualifications

  • Experience deploying AI workloads on Kubernetes.

  • Experience with NVIDIA GPU Operator.

  • Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).

  • Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.

  • Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.

  • Familiarity with MLPerf benchmarking.

  • Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.

  • Experience automating infrastructure using Ansible, Terraform, or similar tools.

  • Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.

Skills Required

  • 7+ years designing or operating large-scale Linux infrastructure
  • 5+ years supporting production GPU clusters for AI or HPC workloads
  • Experience building multi-node GPU training environments from the ground up
  • Deep expertise with distributed PyTorch training
  • Extensive experience troubleshooting and optimizing NCCL communications
  • Understanding of distributed AI communication patterns, including AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications
  • Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems
  • MLPerf benchmarking experience
  • Understanding of GPU memory management, including KV cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism
  • Experience optimizing LLM inference throughput, including tokens-per-second optimization, batch sizing, continuous batching, KV cache tuning, and memory bandwidth optimization
  • Experience tuning CUDA, NCCL, UCX, and MPI for distributed performance
  • Expert-level Linux systems administration skills
  • Experience with Slurm workload manager
  • Experience using Pyxis and Enroot for containerized GPU workloads
  • Strong Python and Bash scripting skills
  • Experience deploying AI workloads on Kubernetes
  • Experience with NVIDIA GPU Operator
  • Experience with Kubernetes batch scheduling tools such as Volcano, Kueue, or Run:ai
  • Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang
  • Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments
  • Familiarity with MLPerf benchmarking
  • Experience deploying Prometheus, Grafana, and DCGM Exporter monitoring solutions
  • Experience automating infrastructure with Ansible, Terraform, or similar tools
  • Experience working in AWS, Azure, or GCP GPU environments in addition to bare metal
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
99 Employees
Year Founded: 2016

What We Do

STN, Inc. is a managed technology and infrastructure provider serving enterprise, regulated, and AI-driven organizations. It designs, operates, and supports secure, scalable systems, including managed IT, cloud and platform services, cybersecurity, data management, compliance engineering, enterprise hardware and software, and GPU One, its GPU-as-a-Service platform for AI training, tuning, inference, and other high-performance workloads. STN emphasizes reliability, audit readiness, and ongoing human support.

Similar Jobs

NVIDIA Logo NVIDIA

Distinguished Engineer, Production Engineering, Cluster Management

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Remote or Hybrid
2 Locations
21960 Employees
320K-489K Annually

NVIDIA Logo NVIDIA

Senior HPC AI Cluster Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees
176K-334K Annually

Backblaze Logo Backblaze

Cluster & Systems Capacity Engineer

Cloud • Information Technology
Remote
United States
363 Employees
123K-175K Annually

Similar Companies Hiring

Rain Thumbnail
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
New York, NY
100 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account