Senior Infrastructure Engineer

Posted 19 Days Ago
Be an Early Applicant
2 Locations
In-Office
200K-330K Annually
Senior level
Artificial Intelligence • On-Demand • Software
The AI Infrastructure Platform: scalable, efficient, on-demand GPUs
The Role
Design and scale infrastructure powering a global GPU marketplace. Responsibilities include GPU provisioning, workload scheduling, orchestration, provider onboarding, usage tracking, billing, pricing logic, resource management, and marketplace APIs. The engineer will optimize performance, reliability, security, fault tolerance, and multi-tenant cloud systems while collaborating with product and infrastructure teams. The role requires strong Python and C++ development, distributed systems expertise, and experience building large-scale compute platforms.
Summary Generated by Built In
About Us

Vast.ai’s cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing—reshaping our future for the benefit of humanity.

We are a growing and highly motivated team dedicated to an ambitious technical plan. Our structure is flat, our ambitions are out‑sized, and leadership is earned by shipping excellence.

We seek engineers with strong intrinsic drive, a true passion for advancing the state of the art, and a mix of architecture, coding, and communication skills.

LOCATION: On-site at our office in San Francisco or Westwood, Los Angeles.

About the Role

As a Senior Infrastructure Engineer, you will help design and scale the core systems that power Vast.ai’s global GPU marketplace.

You’ll work closely with our founders and core engineering team to extend the underlying compute infrastructure — from GPU provisioning and scheduling to billing, orchestration, and marketplace dynamics.

We’re looking for someone who has previously built large-scale infrastructure platforms — systems with similarities to Vast.ai, or distributed compute orchestration frameworks.

Full-time · On-site at either our SF or LA offices

Tech Stack

Python, C++, PostgreSQL, Linux, Docker, KVM, Redis, Terraform, AWS, REST/gRPC APIs

Ideal Experience
  • Distributed Systems: Experience building high-throughput backend systems or compute clouds

  • Compute Orchestration: Familiarity with Docker, or custom scheduling frameworks

  • GPU Infrastructure: Understanding of GPU provisioning, driver management, and workload scheduling

  • Billing & Metering: Implemented or integrated usage-based billing and account credit systems

  • Marketplace Dynamics: Knowledge of dynamic pricing, spot instances, or supply-demand balancing mechanisms

  • Security & Multi-Tenancy: Experience designing secure, multi-tenant systems in cloud environments

  • Programming: Strong programming skills in Python and C++; ability to write performant, maintainable, well-architected code

  • Database Expertise: Comfortable designing schemas and queries for large-scale data systems (PostgreSQL preferred)

Bonus points for:

  • Experience with GPU security, virtualization, or zero-trust compute isolation

  • Prior startup experience or end-to-end product ownership

Key Responsibilities
  • Improve the backend systems that power Vast.ai’s compute marketplace

  • Integrate GPU provider onboarding, usage tracking, billing, and orchestration APIs

  • Develop scalable infrastructure for workload scheduling and resource management

  • Optimize pricing and marketplace logic for efficiency and transparency

  • Benchmark, profile, and harden systems for performance, reliability, and fault tolerance

  • Collaborate with product and infrastructure teams to shape the future of decentralized compute

Interview Process

After submitting your application, our technical team reviews your credentials. If selected, you'll proceed through the following stages:

  • 15 min - Initial screening (virtual)

  • 45 min - Quick dive into Vast, work history (virtual)

  • 45 min - Systems and architectures (virtual)

  • 1 hour - LLM-assisted coding assessment (virtual)

  • 2 hours - Meet and greet with coding assessment (on-site)

Our goal is to complete the interview process in two weeks.

Benefits
  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

Skills Required

  • Experience building high-throughput backend systems, compute clouds, or distributed systems
  • Familiarity with Docker or custom compute scheduling and orchestration frameworks
  • Understanding of GPU provisioning, driver management, and workload scheduling
  • Experience implementing or integrating usage-based billing and account credit systems
  • Knowledge of dynamic pricing, spot instances, or supply-demand balancing mechanisms
  • Experience designing secure, multi-tenant systems in cloud environments
  • Strong programming skills in Python and C++
  • Ability to write performant, maintainable, well-architected code
  • Experience designing schemas and queries for large-scale data systems, preferably PostgreSQL
  • Experience with GPU security, virtualization, or zero-trust compute isolation
  • Prior startup experience or end-to-end product ownership
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Los Angeles, CA
41 Employees
Year Founded: 2018

What We Do

Vast.ai is the market leader for low cost GPU rentals. The service connects data centers and professionals running the Vast hosting software with users who can quickly find the best deals for compute according to their specific requirements. Vast.ai GPU rentals are ~3-5X cheaper than current alternatives. Consumer computers and consumer GPUs in particular are considerably more cost effective than equivalent enterprise hardware. We are helping the millions of underutilized consumer GPUs around the world enter the cloud computing market for the first time.

Similar Jobs

Hybrid
5 Locations
289097 Employees

ServiceNow Logo ServiceNow

Infrastructure Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Mountain View, CA, USA
29000 Employees
191K-334K Annually

CoreWeave Logo CoreWeave

Senior Software Engineer

Cloud • Information Technology • Machine Learning
In-Office
2 Locations
1450 Employees
153K-204K Annually

Headway Logo Headway

Senior Security Engineer

Consumer Web • Healthtech • Professional Services • Social Impact • Software
In-Office or Remote
3 Locations
819 Employees
223K-279K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account