Senior Software Engineer - HPC Cost Optimization & Efficiency

Reposted 17 Days Ago
Be an Early Applicant
Foster City, CA, USA
Hybrid
219K-263K Annually
Senior level
Artificial Intelligence • Machine Learning • Robotics • Software • Transportation • Design • Manufacturing
Zoox is an autonomous mobility company that’s created a purpose-built robotaxi to give the world a better way to ride.
The Role
Drive cost optimization across Zoox's HPC infrastructure to enhance efficiency. Collaborate with teams, implement cost reduction strategies, and modernize cloud platforms.
Summary Generated by Built In
Zoox is looking for an experienced Software Engineer to drive cost optimization and efficiency improvements across our custom High-Performance Computing infrastructure managing  annual compute spend. As Zoox scales its autonomous vehicle development, intelligent resource management and cost efficiency have become critical to our success. You will modernize our HPC platform—built on industry-leading technologies like Ray.io, SLURM, and Kubernetes—with a focus on maximizing utilization, eliminating waste, and reducing cloud costs while maintaining world-class developer velocity.
 
These HPC services form the backbone of development workflows across all Zoox software teams, from data engineering to training our AI models in Perception, Planner, Prediction, to Simulation, and more. You will have a direct impact on Zoox's bottom line through measurable cost reductions and efficiency gains.
 
The position comes with a high degree of independence and the opportunity to define Zoox's compute economics strategy, both technically and organizationally. You will work closely with stakeholders in Autonomy and Software teams to balance performance requirements with cost constraints, incorporating FinOps best practices and the latest cost optimization techniques.

In this role, you will:

  • Design and implement cost optimization strategies across distributed compute infrastructure, targeting millions in annual savings
  • Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap to optimize compute efficiency across all Zoox workloads
  • Create production-grade APIs, SDKs, and tools that make cost-efficient patterns the default developer experience
  • Optimize job scheduling algorithms and auto-scaling policies to maximize resource utilization and minimize idle capacity
  • Design multi-region orchestration strategies that optimize for data locality, cost and performance
  • Identify and eliminate inefficient workload patterns through profiling, analysis, and developer education by coordinating with workload owners across multiple teams
  • Evaluate new technologies and paradigms that reduce cost while meeting Zoox's computational and storage needs
  • Develop cost and demand forecasting models and budget management tools for capacity planning

Qualifications

  • Experience optimizing large-scale distributed systems for cost and efficiency
  • Experience with Ray.io, particularly Ray Core and Ray Data
  • Experience with Kubernetes, particularly for heterogeneous workloads and cost optimization
  • Experience with cloud cost management on AWS (Cost Explorer) or similar providers
  • Track record of achieving measurable cost reductions in production infrastructure
  • Demonstrated ability to prioritize development work and build cross-functional consensus around cost/performance tradeoffs
  • Proficiency with Python

Bonus Qualifications

  • Understanding of FinOps principles and practices
  • Experience building cost attribution, chargeback, or showback systems
  • Exposure to machine learning workloads (training, inference, data generation) from a cost optimization perspective
  • Experience with Kubernetes or SLURM at scale (>10k+ nodes)
  • Experience with SLURM workload manager and advanced scheduling policies
  • Background in algorithmic optimization or operations research

About Zoox
 
Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We’re looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team.
 
Follow us on LinkedIn
 
A Final Note: You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.

Skills Required

  • Experience optimizing large-scale distributed systems for cost and efficiency
  • Experience with Ray.io, particularly Ray Core and Ray Data
  • Experience with Kubernetes, particularly for heterogeneous workloads and cost optimization
  • Experience with cloud cost management on AWS (Cost Explorer) or similar providers
  • Track record of achieving measurable cost reductions in production infrastructure
  • Demonstrated ability to prioritize development work and build cross-functional consensus around cost/performance tradeoffs
  • Proficiency with Python

Zoox Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Zoox and has not been reviewed or approved by Zoox.

  • Healthcare Strength Healthcare is extensive, with broad medical and vision options, company‑paid disability coverage, and multiple mental‑health resources. Feedback suggests coverage breadth and auxiliary programs support a wide range of needs.
  • Parental & Family Support Family supports include paid parental leave, additional pregnancy disability time, fertility coverage, and adoption/surrogacy assistance. Backup care and family‑oriented programs further reinforce support across life stages.
  • Wellbeing & Lifestyle Benefits Day‑to‑day perks are robust, including free daily meals, fitness subsidies, commuter support, and on‑site amenities. Feedback suggests these lifestyle benefits enhance convenience and workplace experience, especially for office‑based roles.

Zoox Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Foster City, CA
2,900 Employees
Year Founded: 2014

What We Do

Zoox is an autonomous mobility company that was founded to provide a safer, cleaner, and more enjoyable future on the road. To achieve that goal, the company has spent the past 10 years creating a purpose-built robotaxi that gives the world a better way to ride.

Why Work With Us

At Zoox, we are working to solve one of the greatest technological challenges of our generation. From the beginning, we have been focused on our goal of reimagining transportation from the ground up. We are a mission-driven community of innovators working together to create a safer, cleaner, and more enjoyable future on the road.

Gallery

Gallery

Similar Jobs

Dynatrace Logo Dynatrace

Solutions Engineer

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
San Francisco, CA, USA
5200 Employees
130K-190K Annually

Airwallex Logo Airwallex

Executive Assistant

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Remote or Hybrid
San Francisco, CA, USA
2000 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate II

eCommerce • Fashion • Other • Retail • Sales • Wearables • Design
Remote or Hybrid
South Coast, CA, USA
16000 Employees
15-24 Hourly

Zeta Global Logo Zeta Global

Lead Software Engineer

AdTech • Artificial Intelligence • Marketing Tech • Software • Analytics
Easy Apply
Hybrid
San Francisco, CA, USA
2429 Employees
150K-200K Annually

Similar Companies Hiring

Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees
Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account