Head of Infrastructure

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
Hybrid
Senior level
Artificial Intelligence • Cloud • Semiconductor • Infrastructure as a Service (IaaS)
The Role
Own the infrastructure layer of an AI inference cloud, including the Kubernetes-based control plane, gateway, load balancing, observability, capacity planning, accelerator operations, and incident response. The role begins hands-on and expands to building and managing the infrastructure team. Responsibilities include supporting heterogeneous ASIC and GPU fleets, partnering with hardware vendors, optimizing tail latency, establishing on-call practices, and bringing up new hardware platforms.
Summary Generated by Built In
About Us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.

About the role

You'll own the infrastructure layer of our inference cloud end-to-end. Today that means the control plane, the gateway in front of our ASIC fleet, and the observability stack that tells us where every millisecond goes. Over the next 6-8 months, it will grow into a heterogeneous fleet: ASICs for decode, GPUs for pre-fill, and the physical-layer ownership that comes with it.

The first six months are hands-on: k8s manifests, dashboards, oncall, and a direct line to our ASIC partner's engineering team when production behaves strangely. The team grows under you from there.

What you'll do:
  • Own the inference control plane. At the moment, it's built on configuration provided by our ASIC partner; you'll be the person who understands it deeply enough to modify, extend, and eventually replace pieces of it.

  • Own the gateway and load balancer that fronts the fleet. Model placement, request routing, and tail-latency engineering live here, driven by live utilization and per-model SLOs.

  • Own observability end-to-end. Per-request tracing from OpenRouter ingress through to the accelerator, with p50/p95/p99 dashboards, SLOs, and alerting that wakes the right person.

  • Run capacity planning against a real, distributed traffic mix across the open-weight models we serve.

  • Own the operational side of the ASIC partnership. Most weird production issues route through their engineering team until we build that expertise in-house, and you'll be our technical face in those conversations.

  • Bring up the pre-fill side of our disaggregated architecture on a second hardware platform as it comes online. Different vendor, different fabric, different kernels.

  • Build the on-call and incident response practice from zero. Hire and grow the team underneath you.

What we need from you:
  • 7+ years in infrastructure, SRE, or platform engineering, with at least some of it at a serious inference, ML, or HPC shop.

  • Hands-on with Kubernetes at production scale — not just deploying, but debugging the weird stuff.

  • Strong instincts for tail latency. You think about p99 and utilization as the same problem, not different ones.

  • Comfortable owning a vendor relationship where the vendor's bugs are now your production issues.

  • Track record of building observability practices that actually catch problems, not just generate dashboards.

  • Have been on-call through real incidents and can talk about what you learned.

  • Want to be the first infra hire at something early, not the tenth at something big.

NIce-to-Haves:
  • Experience operating non-NVIDIA accelerators in production — TPUs, ASICs, or alternative GPU vendors.

  • Background with model-serving stacks (vLLM, TGI, TensorRT-LLM, SGLang).

  • Network fabric experience at data-center scale (RoCE, InfiniBand).

  • Have hired and managed an infra team before.

  • Comfort at the hardware boundary — firmware, drivers, thermals — for when the roadmap takes us there.

Skills Required

  • 7+ years of experience in infrastructure, SRE, or platform engineering
  • Experience at a serious inference, machine learning, or high-performance computing organization
  • Production-scale Kubernetes experience, including debugging complex issues
  • Strong understanding of tail latency, p99 performance, and utilization
  • Experience owning vendor relationships and resolving vendor-related production issues
  • Experience building effective observability practices
  • Experience participating in on-call rotations and handling real incidents
  • Experience operating non-NVIDIA accelerators such as TPUs, ASICs, or alternative GPUs
  • Experience with model-serving stacks such as vLLM, TGI, TensorRT-LLM, or SGLang
  • Data-center-scale network fabric experience with RoCE or InfiniBand
  • Experience hiring and managing an infrastructure team
  • Experience with firmware, drivers, or hardware thermals
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
8 Employees
Year Founded: 2026

What We Do

General Compute is an AI inference cloud company, positioning itself as the neocloud for alternative chips. The company productionizes purpose-built AI silicon from vendors such as SambaNova, Cerebras, Positron, and d-Matrix by purchasing the racks, securing data center space, and operating the hardware for customers. Its architecture splits prefill on GPUs and decode on specialized accelerators, generating tokens 5-7x faster than GPU-based competitors, serving frontier labs, fast-growing AI application companies, and asset-light clouds.

Similar Jobs

NVIDIA Logo NVIDIA

Head of Infrastructure Security Engineering - EDA Clusters

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
5 Locations
21960 Employees
272K-489K Annually

NVIDIA Logo NVIDIA

Head of Global Customer Engineering - Internal EDA Infrastructure

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
5 Locations
21960 Employees
272K-489K Annually

Fuse (f.energy) Logo Fuse (f.energy)

Head of Infrastructure

Energy • Renewable Energy
In-Office
San Leandro, CA, USA
38 Employees
160K-250K Annually

Woven by Toyota Logo Woven by Toyota

Head of Infrastructure Engineering

Automotive • Software • Automation
Hybrid
Palo Alto, CA, USA
1679 Employees
161K-265K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account