AI, HPC & GPU Infrastructure Support Engineer

Posted Yesterday
Be an Early Applicant
Los Angeles, CA, USA
In-Office
90K-150K Annually
Entry level
Artificial Intelligence • On-Demand • Software
The AI Infrastructure Platform: scalable, efficient, on-demand GPUs
The Role
Troubleshoot complex Linux, NVIDIA GPU, CUDA, Docker, virtualization, networking, hardware, and AI workload issues. Own escalated support tickets, assist clients and infrastructure suppliers, identify root causes, and collaborate with engineering teams. Build Python and Bash diagnostic tooling, create runbooks, improve monitoring and troubleshooting processes, and support supplier onboarding and machine management. The role is fully or mostly on-site in Los Angeles, with schedules covering Monday-Friday or Sunday-Thursday.
Summary Generated by Built In

About Us

Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This role focuses on troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, GPU workloads, Ubuntu, Docker, KVM based virtual machines, networking, hardware, BIOS, and firmware. You’ll investigate failures, reproduce issues, identify root causes, and propose practical solutions across the full infrastructure stack.

You’ll also serve as the engineering resource our L1 support team relies on when tickets go beyond frontline triage. You’ll own complex escalations end-to-end, gather technical evidence, coordinate with the appropriate teams, and communicate findings clearly to clients, infrastructure suppliers, and internal teams.

The best engineers in this role don’t just resolve individual issues—they recognize recurring patterns, improve diagnostic tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic Linux, GPU, and infrastructure problems.

Strong GPU troubleshooting experience, Linux systems knowledge, and technical support skills are the primary requirements. You should be comfortable working autonomously in Ubuntu environments and troubleshooting NVIDIA drivers, CUDA, containers, virtual machines, networking, hardware, and GPU workloads.

Vast.ai users or hosts strongly preferred.

Location and Schedule

This is a full-time position based in our Westwood, Los Angeles office.

Available schedules:

  • Monday–Friday: Fully on-site

  • Sunday–Thursday: Four days on-site and one day working from home

Key Responsibilities

  • Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments

  • Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks

  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads

  • Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

  • Provide managed support for supplier onboarding and ongoing machine management, including installation, configuration, and post setup troubleshooting

  • Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance

  • Provide coverage for L1 support overflow during peak periods or incidents

  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

  • Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues

You Are

  • Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line

  • Someone who enjoys debugging difficult problems and fixing broken systems

  • Methodical and focused on finding root causes, not just temporary fixes

  • Able to manage complex tickets independently

  • A clear written communicator with an interest in AI infrastructure and GPU computing

Must-Haves

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions

  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting

  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting

  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting

  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting

  • Python and Bash scripting skills for automation and diagnostic tooling

  • Strong written English communication that is clear, professional, and technically precise

  • Experience providing technical support in a customer-facing or internal help desk environment

  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice-to-Haves

  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers

  • Monitoring and observability experience (Prometheus, Grafana)

  • Relevant certifications: RHCSA, CompTIA Linux+, or similar

  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Interview Process (~1 week)

After you submit your application, our technical team will review your experience and qualifications. Selected candidates will proceed through the following stages:

  • 15 minutes — Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role

  • 45 minutes — Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience

  • 2 hours — Meet and Greet and Technical Assessment (On-site): Meet the team and complete an LLM-assisted Linux systems operations assessment

Annual Salary Range

$90,000 – $160,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits
  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

 

Skills Required

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions
  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting
  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting
  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and Layer 2/Layer 3 troubleshooting
  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting
  • Python and Bash scripting skills for automation and diagnostic tooling
  • Strong written English communication that is clear, professional, and technically precise
  • Experience providing technical support in a customer-facing or internal help desk environment
  • Ability to prioritize concurrent escalated tickets by severity and customer impact, balancing reactive resolution with documentation and tooling work
  • Familiarity with AI/ML frameworks such as TensorFlow and PyTorch and running GPU-accelerated containers
  • Monitoring and observability experience with Prometheus and Grafana
  • Relevant certifications such as RHCSA, CompTIA Linux+, or similar
  • Knowledge of the Vast.ai platform as a client or infrastructure supplier
  • Vast.ai user or infrastructure supplier experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Los Angeles, CA
41 Employees
Year Founded: 2018

What We Do

Vast.ai is the market leader for low cost GPU rentals. The service connects data centers and professionals running the Vast hosting software with users who can quickly find the best deals for compute according to their specific requirements. Vast.ai GPU rentals are ~3-5X cheaper than current alternatives. Consumer computers and consumer GPUs in particular are considerably more cost effective than equivalent enterprise hardware. We are helping the millions of underutilized consumer GPUs around the world enter the cloud computing market for the first time.

Similar Jobs

PwC Logo PwC

Pricing and Revenue Consulting Manager - Consumer Markets Sector

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
58 Locations
370000 Employees
99K-232K Annually

PwC Logo PwC

Sanctions - Crypto & Digital Assets - Senior Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
8 Locations
370000 Employees
124K-280K Annually

PwC Logo PwC

Assurance Internal Communications Director

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
69 Locations
370000 Employees
123K-123K Annually

PwC Logo PwC

Consultant

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
8 Locations
370000 Employees
99K-232K Annually

Similar Companies Hiring

Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account