ML Platform Engineer

Posted 12 Days Ago
San Francisco, CA, USA
In-Office
Entry level
Artificial Intelligence • Information Technology • Automation
The Role
Build and operate the ML execution platform supporting data generation, simulation orchestration, training, fine-tuning, benchmarking, and production deployments. Manage scalable compute, storage, networking, observability, reproducibility, CI/CD, and customer deployments across cloud, hybrid, and on-premises environments. Partner directly with customers and the founding team to deliver reliable, secure, and repeatable systems for industrial AI workloads.
Summary Generated by Built In

📍 San Francisco | 🏢 5 Days Onsite

Location: Onsite in San Francisco

Compensation: Competitive Salary + Equity

 

Who We Are

Engineering simulation is one of the last major categories of software that AI hasn't rebuilt. The tools used to design aircraft, ships, reservoirs, and medical devices still run on numerical methods that are decades old, and an engineer can wait a full day for a single answer. UniversalAGI is building foundation models that learn physics directly from data, and they are already running in early deployments on real computational fluid dynamics and reservoir engineering problems for some of the largest industrial and defense organizations in the world.

We are a team of 25 researchers and engineers in San Francisco backed by Elad Gil (#1 Solo VC), Eric Schmidt (former Google CEO), Prith Banerjee (ANSYS CTO), Ion Stoica (Databricks Founder), Jared Kushner (former Senior Advisor to the President), David Patterson (Turing Award Winner), and Luis Videgaray (former Foreign and Finance Minister of Mexico).

About the Role

UniversalAGI is hiring a ML Platform Engineer to build and own the execution platform powering our research and customer deployments: data generation + simulation orchestration + training/fine-tuning infrastructure + benchmarking pipelines + production deployments in customer environments.

You’ll work closely with the CEO and founding team to turn research into repeatable, scalable, reliable systems - internally and in customer infrastructure. This is a “ship outcomes” role: your work directly determines how fast we can iterate, how reproducible our results are, and how reliably we deliver in production.

What You’ll Do


Build the foundation platform (internal):

  • Build and operate scalable infrastructure for data generation and simulation workflows (job orchestration, scheduling, queues, retries, observability).

  • Build reproducible pipelines for training/fine-tuning and benchmarking (artifact/version management, experiment tracking, dataset lineage).

  • Own cost/performance tradeoffs across compute, storage, networking, and runtime efficiency.

Deploy to customers (external):

  • Lead deployments of our stack into customer cloud/on-prem environments (AWS/GCP/Azure + hybrid), including secure networking, permissions, and data movement.

  • Build robust deployment patterns: environment provisioning, CI/CD, rollbacks, monitoring, and incident response.

  • Partner with customers to ensure reliability and repeatability under real-world constraints (security, compliance, infra limits, data governance).

Qualifications

  • Strong software engineering skills (clean code, debugging, reliability, reproducibility).

  • Hands-on experience building/operating infrastructure for ML/compute-heavy workflows: pipelines, job orchestration, GPU compute, storage, CI/CD, monitoring.

  • Olympic athlete mindset: You have high standards for yourself and are obsessed with measurable improvement on the metrics you are delivering to customers.

  • Resourcefulness: you know when to do the “quick & correct” fix vs. when to invest in a robust solution, and you can justify the tradeoff with impact/

  • Ownership: Comfortable owning work end-to-end and being accountable for measurable outcomes.

Bonus Qualifications

  • Experience with workflow orchestration (e.g., Ray, Kubernetes, Slurm).

  • Experience with GPU infrastructure and distributed training systems.

  • Experience building evaluation/benchmarking frameworks with strong reproducibility guarantees.

  • Experience deploying into regulated / security-sensitive environments (gov/defense/enterprise).

  • Experience with simulation/HPC pipelines (CFD, meshing, batch workloads) is a plus but not required.

  • Experience in an FDE-style / delivery execution role (or similar “ship results fast” environments).

Cultural Fit

  • Technical Respect: Ability to earn respect through hands-on technical contribution

  • Intensity: Thrives in our unusually intense culture - willing to grind when needed

  • Customer Obsession: Passionate about solving real customer problems, not just cool tech

  • Deep Work: Values long, uninterrupted periods of focused work over meetings

  • High Availability: Ready to be deeply involved whenever critical issues arise

  • Communication: Can translate complex technical concepts to customers and team

  • Growth Mindset: Embraces the compounding returns of intelligence and continuous learning

  • Startup Mindset: Comfortable with ambiguity, rapid change, and wearing multiple hats

  • Work Ethic: Willing to put in the extra hours when needed to hit critical milestones

  • Team Player: Collaborative approach with low ego and high accountability

 

What We Offer

  • Opportunity to shape the technical foundation of a rapidly growing foundational AI company.

  • Work on cutting-edge industrial AI problems with immediate real-world impact.

  • Direct collaboration with the founder & CEO and ability to influence company strategy

  • Competitive compensation with significant equity upside.

  • In-person first culture - 5 days a week in office with a team that values face-to-face collaboration.

  • Access to world-class investors and advisors in the AI space.

 

Benefits

We provide great benefits, including:

  • Competitive compensation and equity.

  • Competitive health, dental, vision benefits paid by the company.

  • 401(k) plan offering.

  • Flexible vacation.

  • Team Building & Fun Activities.

  • Great scope, ownership and impact.

  • AI tools stipend.

  • Monthly commute stipend.

  • Monthly wellness / fitness stipend.

  • Daily office lunch & dinner covered by the company.

  • Immigration support.

 

How We’re Different

 

“The credit belongs to the man who is actually in the arena, whose face is marred by dust and sweat and blood; who strives valiantly; who errs, who comes short again and again... who at the best knows in the end the triumph of high achievement, and who at the worst, if he fails, at least fails while daring greatly." - Teddy Roosevelt

 

At our core, we believe in being “in the arena.” We are builders, problem solvers, and risk-takers who show up every day ready to put in the work: to sweat, to struggle, and to push past our limits. We know that real progress comes with missteps, iteration, and resilience. We embrace that journey fully knowing that daring greatly is the only way to create something truly meaningful.

 

If you're ready to join the future of physics simulation, push creative boundaries, and deliver impact, UniversalAGI is the place for you.

Skills Required

  • Strong software engineering skills, including clean coding, debugging, reliability, and reproducibility
  • Hands-on experience building or operating infrastructure for ML or compute-heavy workflows
  • Experience with pipelines, job orchestration, GPU compute, storage, CI/CD, and monitoring
  • Experience with workflow orchestration tools such as Ray, Kubernetes, or Slurm
  • Experience with GPU infrastructure and distributed training systems
  • Experience building reproducible evaluation and benchmarking frameworks
  • Experience deploying into regulated, security-sensitive, government, defense, or enterprise environments
  • Experience with simulation or HPC pipelines, including CFD, meshing, or batch workloads
  • Experience in an FDE-style delivery execution role or similar environment
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
4 Employees
Year Founded: 2025

What We Do

UniversalAGI is automating physical systems engineering across the entire product lifecycle with artificial intelligence.

Similar Jobs

Braze Logo Braze

Machine Learning Engineer

Marketing Tech • Mobile • Software
Easy Apply
Hybrid
San Francisco, CA, USA
2000 Employees
184K-348K Annually

Physical Intelligence Logo Physical Intelligence

Machine Learning Engineer

Artificial Intelligence • Machine Learning • Robotics
In-Office
San Francisco, CA, USA
191 Employees

Physical Intelligence Logo Physical Intelligence

Machine Learning Engineer

Artificial Intelligence • Machine Learning • Robotics
In-Office
San Francisco, CA, USA
191 Employees

Fetch Logo Fetch

Senior Machine Learning Engineer

AdTech • Big Data • Marketing Tech • Mobile • Grocery & Supermarkets • Consumer Packaged Goods (CPG)
In-Office or Remote
7 Locations
950 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account