Senior AI Infrastructure Engineer — HPC & Compute Clusters

Posted Yesterday
3 Locations
Remote or Hybrid
Senior level
Artificial Intelligence • Computer Vision • Machine Learning • Robotics
The Role
Design, deploy, and maintain large-scale bare-metal GPU clusters for distributed training. Tune Slurm schedulers and Kubernetes integrations, optimize storage and high-speed networking, and build observability pipelines to ensure reliable, high-throughput GPU workloads for distributed deep learning.
Summary Generated by Built In
About US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • GPU Cluster Orchestration: Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes to power multi-node distributed training runs.

  • Workload Management & Scheduler Tuning: Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling for optimal GPU utilization and job throughput.

  • Kubernetes Integration: Manage K8s clusters and hybrid Slurm-K8s environments (e.g., KubeRay, MPI Operator, Slurm on K8s) for serving, data processing pipelines, and interactive research environments.

  • Storage & Network Performance: Optimize high-speed interconnects (InfiniBand/RoCE, NCCL, NVLink) and high-throughput distributed storage (Lustre, WEKA, Ceph) to keep tens of thousands of GPUs continuously saturated.

  • Reliability & Monitoring: Build observability pipelines (Prometheus, Grafana, DCGM) to proactively detect hardware degradation, GPU silent errors, network flaps, and node failures before they impact training jobs.

Requirements
  • You have a Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering.

  • You have deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration).

  • You possess strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems (e.g., NFS, NVMe-oF, or distributed filesystems).

  • You are proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting).

  • Background supporting distributed deep learning frameworks (DeepSpeed, Ray).

Nice to Have
  • Experience managing high-density GPU infrastructure (NVIDIA H100/B200/B300 systems, DGX/HGX architectures, DCGM, NCCL tuning).

  • Experience with high-speed network fabrics including InfiniBand (Subnet Manager, SMDB) and RoCE (v2).

Skills Required

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on HPC/infrastructure experience.
  • Deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share, topology).
  • Strong troubleshooting skills in low-level Linux networking, kernel tuning, hardware diagnostics, and storage systems (NFS, NVMe-oF, distributed filesystems).
  • Proficient with automation and infrastructure-as-code tools (Ansible, Terraform, Helm) and scripting (Python, Bash).
  • Background supporting distributed deep learning frameworks (DeepSpeed, Ray).
  • Experience managing high-density GPU infrastructure (NVIDIA H100/B200/B300, DGX/HGX, DCGM, NCCL tuning).
  • Experience with high-speed network fabrics including InfiniBand (Subnet Manager, SMDB) and RoCE v2.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Veeda AI is a small, fast-moving team of engineers and researchers building the next generation of multimodal foundation world models for Physical AI. Its work sits at the intersection of artificial intelligence, robotics, and embodied intelligence, with engineering roles involving high-throughput image and video data pipelines. The company aims to advance intelligent systems capable of operating in and understanding the physical world.

Similar Jobs

Rubrik Logo Rubrik

Senior Sales Engineer

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Remote
Switzerland
3000 Employees

Drata Logo Drata

Enterprise Account Executive

Security • Software • Cybersecurity • Automation
Remote
26 Locations
600 Employees
194K-273K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Director Sales

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
4 Locations
4600 Employees

Zapier Logo Zapier

Back-end Engineer

Artificial Intelligence • Productivity • Software • Automation
Remote
32 Locations
800 Employees
211K-316K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account