Software Engineer/Senior Software Engineer, Data & ML Platform

Posted Yesterday
Santa Clara, CA, USA
Hybrid
135K-200K Annually
Senior level
Software
The Role
Build and operate production Kubernetes-based compute platform for large-scale data and ML workloads. Improve cluster reliability, GitOps delivery, multi-tenant scheduling, resource isolation, GPU workload orchestration, and reusable batch/workflow systems for petabyte-scale processing while contributing to QMS and continuous improvement.
Summary Generated by Built In

All key offline workloads — large-scale data processing, simulation, auto-labeling, scenario mining, and model training — run on the compute platform this role owns. In this role, you will improve the reliability and efficiency of our Kubernetes infrastructure, make workload onboarding simpler and more self-service, and build reusable batch and workflow capabilities for petabyte-scale processing. We are looking for strong Kubernetes and platform-engineering fundamentals, depth in at least one adjacent area—distributed data processing, ML/GPU infrastructure, or multi-tenant compute systems—and the curiosity and ownership to grow across the others.


We are open to candidates at either the Software Engineer or Senior Software Engineer level. Level will be determined by experience, technical depth, scope of ownership, and demonstrated impact. You do not need experience with every technology in our stack; we value strong fundamentals, ownership, and the ability to learn.


Responsibilities:
  • Operate and evolve our production Kubernetes clusters end to end: bare-metal provisioning automation, highly available control planes, node lifecycle, GPU container runtime, networking, and storage

  • Build safe, repeatable GitOps-based delivery for platform services and user applications using tools such as Argo CD, Helm, and Kustomize

  • Develop shared multi-tenant platform capabilities for scheduling, resource isolation, storage, networking, access control, secrets, and observability while improving CPU/GPU utilization and cost efficiency

  • Build and improve reusable distributed batch and workflow platforms for Spark data processing and GPU-based replay and simulation

  • Ensure that your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous improvement efforts

Required Skills:

  • BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience

  • Hands-on experience operating production Kubernetes clusters — node lifecycle, upgrades, troubleshooting — plus GitOps and infrastructure-as-code experience

  • Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management

  • Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end

Preferred Skills:

  • Experience with Ray or Kubeflow

  • Experience with lakehouse technologies such as Delta Lake or Apache Iceberg

  • Experience operating large-scale distributed data-processing and workflow systems, with hands-on depth in a system such as Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator

Skills Required

  • BS, MS, or PhD in Computer Science or related technical field, or equivalent practical experience
  • Hands-on experience operating production Kubernetes clusters (node lifecycle, upgrades, troubleshooting) plus GitOps and infrastructure-as-code experience
  • Experience with GPU or ML workload scheduling, queueing, priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management
  • Strong Kubernetes and platform-engineering fundamentals and depth in at least one adjacent area (distributed data processing, ML/GPU infrastructure, or multi-tenant compute systems)
  • Self-driven with strong sense of ownership; able to drive projects end-to-end
  • Experience with Ray or Kubeflow
  • Experience with lakehouse technologies such as Delta Lake or Apache Iceberg
  • Experience operating large-scale distributed data-processing and workflow systems, e.g., Apache Spark and Argo Workflows

Plus Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Plus and has not been reviewed or approved by Plus.

  • Leave & Time Off Breadth Unlimited PTO in addition to company holidays and flexible work arrangements are offered, indicating broad time-off flexibility. This setup signals strong support for taking time away from work.
  • Healthcare Strength Tiered medical, dental, and vision options allow employees to select coverage that fits their needs. This breadth of core health coverage aligns with a comprehensive benefits approach.
  • Wellbeing & Lifestyle Benefits Daily catered lunches at key offices and company-sponsored professional development add meaningful day-to-day and growth-oriented perks. These offerings enhance overall wellbeing and workplace experience.

Plus Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Cupertino, CA
471 Employees
Year Founded: 2016

What We Do

Plus is a global provider of highly automated driving and fully autonomous driving solutions. Named by Forbes as one of America's Best Startup Employers and Fast Company as one of the World’s Most Innovative Companies, Plus's customers are already operating its product on the road today. Working with one of the largest companies in the U.S., vehicle manufacturers and others, Plus is making transportation safer and greener. Plus has received a number of industry awards and distinctions for its transformative technology and business momentum from Fast Company, Insider, Consumer Electronics Show, AUVSI, and others. For more information, visit www.plus.ai

Similar Jobs

Axle Health Logo Axle Health

Head of Growth Marketing

Artificial Intelligence • Healthtech • Information Technology • Logistics
In-Office
Santa Monica, CA, USA
25 Employees
170K-200K Annually
Easy Apply
In-Office
Los Angeles, CA, USA
24 Employees
30-45 Hourly
Easy Apply
In-Office
Los Angeles, CA, USA
24 Employees
160K-220K Annually

Legora Logo Legora

Legal Engineer - In-House

Artificial Intelligence • Legal Tech • Software
In-Office
San Francisco, CA, USA
700 Employees
225K-325K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account