GPU Systems Engineer – HPC / Parallel Computing

Posted 3 Hours Ago
Be an Early Applicant
2 Locations
In-Office
160K-320K Annually
Entry level
Artificial Intelligence • On-Demand • Software
The AI Infrastructure Platform: scalable, efficient, on-demand GPUs
The Role
Design and optimize GPU kernels and tensor libraries for scalable AI inference. Apply HPC and parallel-computing techniques, evaluate emerging GPU architectures and resource-management approaches, and improve GPU infrastructure efficiency. The role requires advanced C++ development, parallel programming expertise, systems optimization, and performance tooling experience. This is a full-time, on-site position in San Francisco or Los Angeles.
Summary Generated by Built In
About Us

Vast.ai’s cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing—reshaping our future for the benefit of humanity.

We are a growing and highly motivated team dedicated to an ambitious technical plan. Our structure is flat, our ambitions are out‑sized, and leadership is earned by shipping excellence.

We seek engineers with strong intrinsic drive, a true passion for advancing the state of the art, and a mix of architecture, coding, and communication skills.

LOCATION: On-site at our office in San Francisco or Westwood, Los Angeles.

About the Role

We’re looking for a systems engineer with HPC or parallel programming experience to help scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding edge of AI.

  • Full-Time

  • On-site at either our SF or LA offices

Tech Stack

CUDA/C++, GPGPU, Python, Linux

Key Responsibilities
  • Design and optimize GPU kernels and tensor libraries

  • Translate HPC techniques into scalable AI inference solutions

  • Evaluate emerging architectures and resource management approaches

  • Collaborate with technical leadership to improve GPU infrastructure efficiency

Ideal Experience
  • Advanced C++ (C++17/20 preferred)

  • Expertise with at least one parallel framework (CUDA, HIP, SYCL, OpenCL, OpenACC, or similar)

  • Strong background in systems optimization and HPC performance tooling

  • Familiarity with distributed training/inference frameworks (bonus)

Interview Process

After submitting your application, our technical team reviews your credentials. If selected, you'll proceed through the following stages:

  • 15 min - Initial screening (virtual)

  • 45 min - Quick dive into Vast, work history (virtual)

  • 45 min - Systems and architectures (virtual)

  • 1 hour - LLM-assisted coding assessment (virtual)

  • 2 hours - Meet and greet with coding assessment (on-site)

Our goal is to complete the interview process in two weeks.

Benefits
  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

Skills Required

  • Advanced C++ experience, preferably with C++17 or C++20
  • Expertise with at least one parallel programming framework, such as CUDA, HIP, SYCL, OpenCL, or OpenACC
  • Strong background in systems optimization and HPC performance tooling
  • Familiarity with distributed training or inference frameworks
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Los Angeles, CA
41 Employees
Year Founded: 2018

What We Do

Vast.ai is the market leader for low cost GPU rentals. The service connects data centers and professionals running the Vast hosting software with users who can quickly find the best deals for compute according to their specific requirements. Vast.ai GPU rentals are ~3-5X cheaper than current alternatives. Consumer computers and consumer GPUs in particular are considerably more cost effective than equivalent enterprise hardware. We are helping the millions of underutilized consumer GPUs around the world enter the cloud computing market for the first time.

Similar Jobs

SailPoint Logo SailPoint

Customer Success Manager

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
United States
2461 Employees
125K-210K Annually

MongoDB Logo MongoDB

Senior Solutions Architect

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
San Francisco, CA, USA
5550 Employees
104K-204K Annually

Deepgram Logo Deepgram

Director, Text-to-Speech Synthesis Research

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
3 Locations
150 Employees
213K-328K Annually

CrowdStrike Logo CrowdStrike

Consultant

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
115K-160K Annually

Similar Companies Hiring

Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account