Founding Platform Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
Hybrid
Entry level
Artificial Intelligence • Cloud • Semiconductor • Infrastructure as a Service (IaaS)
The Role
Build the foundational inference cloud platform, including its control plane, API, serving layer, routing, model placement, scheduling, reliability, and observability. The role involves scaling a heterogeneous GPU and ASIC fleet, partnering with data center and model deployment teams, and shaping platform architecture as an early individual contributor at a small startup.
Summary Generated by Built In
About us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.

About the role

You will build the inference cloud itself — the control plane, API, and serving layer that turn racks into a sellable product. There's no existing platform team to inherit or manage, no legacy system to work around, and no established playbook to follow — just the platform itself to build, with reliability treated as core infrastructure from day one rather than something bolted on after the first outage.The technical problem is also genuinely unsolved elsewhere. The fleet is heterogeneous by design — GPUs for prefill, multiple ASIC vendors for decode — so there's no single-vendor playbook to lean on; you'll be defining how a mixed-hardware inference cloud gets scheduled, routed, and served reliably, in close partnership with the teams standing up the physical fleet.

What you'll do:
  • Build/own the control plane — routing, model placement, scheduling across a mixed ASIC/GPU pool

  • Build the API and serving layer exposing rack capacity as a sellable product

  • Build in reliability and observability from day one

  • Scale the platform ahead of the demand curve

  • Partner closely with data center deployment and model bring-up teams

  • Be a founding technical voice on platform architecture

What we need from you:
  • Strong systems engineering background on cloud control planes/serving infra at scale

  • Comfort being a high-impact IC rather than a manager

  • Track record building reliability from scratch

  • Comfort with hardware heterogeneity/ambiguity

  • Genuine interest in being an early hire at a ~6–7 person company.

Nice-to-haves:
  • LLM-serving infra experience (vLLM, TGI, Ray Serve, etc.)

  • Experience running non-NVIDIA accelerators (TPUs/ASICs) in production.

Skills Required

  • Strong systems engineering background with cloud control planes or serving infrastructure at scale
  • Comfort working as a high-impact individual contributor rather than a manager
  • Track record of building reliability from scratch
  • Comfort working with hardware heterogeneity and ambiguity
  • Genuine interest in joining an early-stage company of approximately six to seven people
  • Experience with LLM-serving infrastructure such as vLLM, TGI, or Ray Serve
  • Experience running non-NVIDIA accelerators such as TPUs or ASICs in production
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
8 Employees
Year Founded: 2026

What We Do

General Compute is an AI inference cloud company, positioning itself as the neocloud for alternative chips. The company productionizes purpose-built AI silicon from vendors such as SambaNova, Cerebras, Positron, and d-Matrix by purchasing the racks, securing data center space, and operating the hardware for customers. Its architecture splits prefill on GPUs and decode on specialized accelerators, generating tokens 5-7x faster than GPU-based competitors, serving frontier labs, fast-growing AI application companies, and asset-light clouds.

Similar Jobs

Vincer Logo Vincer

Software Engineer

Artificial Intelligence • Healthtech • Sales • Software
Hybrid
2 Locations
160K-190K Annually

Converge Logo Converge

Platform Engineer

AdTech • Big Data • Marketing Tech • Analytics
In-Office
New York, NY, USA
4 Employees

ServiceNow Logo ServiceNow

Regional Brand Creative Lead

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
New York, NY, USA
29000 Employees
148K-260K Annually

Bilt Logo Bilt

Lead Product Designer

Fintech • Mobile • Real Estate • Financial Services • PropTech
In-Office
New York, NY, USA
200 Employees
160K-215K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account