Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.
Job SummaryWe’re hiring a lead for the data platform: camera on a worker to training-ready datasets, and out to research customers. Collection is monocular 1920×1080p 30fps in the wild, targeting 100M hours. The roles we mean: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, Eventual/Daft. This is not a warehouse, analytics, or generic backend seat.
Key ResponsibilitiesOwn the data platform end-to-end: on-device capture, upload under flaky bandwidth, object storage, training-ready shards
Compression, codecs, and storage-tier trade-offs so 1080p30 hours stay cheap enough to keep collecting
Upload that survives bad networks: on-device buffering, batching, retries, a drop rate you can actually see
Object storage and training-shard formats. The hard problem is petabyte-scale media, not a warehouse
Own dataset packaging, versioning, and delivery to external research customers
Work with Shenzhen firmware so new devices speak one ingest contract, not a custom path per SKU
Make health, cost, and drop rate obvious as we add sites and countries
You have owned a production media or sensor data path at real scale: object storage at petabyte scale, video codecs and compression, upload under flaky bandwidth, or training-shard / dataset formats
That kind of data path: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, or Eventual/Daft. Demo-scale ETL is not this job
Strong software engineering. Python and at least one systems language. Linux
You measure cost and throughput, not whether the demo uploaded
You want to scale in-the-wild physical-labor video, not run a generic data org
Pose, multi-camera, or other large media besides video
Cloud (AWS or GCP), orchestration (Kubernetes, Airflow, Temporal), or IaC
Dataset management or annotation tooling
You have shipped dataset delivery to external research or training customers
Competitive pay
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2k per month for those living within walking distance of the office
Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification
Unlimited Codex and Claude credits
Travel
Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: [email protected]
Skills Required
- Production ownership of a media or sensor data path at real scale, including petabyte-scale object storage, video codecs and compression, unreliable-network uploads, or training-shard and dataset formats.
- Experience with large-scale data paths such as autonomous vehicle, video, media encoding, dataset, or related infrastructure platforms.
- Strong software engineering skills.
- Proficiency in Python and at least one systems programming language.
- Linux experience.
- Ability to measure and optimize cost and throughput.
- Experience with pose, multi-camera, or other large media data.
- Experience with AWS or GCP.
- Experience with Kubernetes, Airflow, or Temporal.
- Infrastructure as code experience.
- Dataset management or annotation tooling experience.
- Experience delivering datasets to external research or training customers.
What We Do
Build AI is a public benefit corporation and data hyperscaler for Physical AI. It integrates hardware manufacturing, logistics, data collection, and model training to scale egocentric physical-labor datasets for researchers and labs. Its mission is to solve physical labor and unlock human potential, advancing robotics and physical superintelligence. The company operates in San Francisco and Shenzhen and develops economically useful human-data infrastructure.









