Staff Software Engineer - Compute

Posted 15 Hours Ago
Be an Early Applicant
3 Locations
Remote or Hybrid
314K-465K Annually
Expert/Leader
Software
The Role
Lead design and implementation of a highly available GPU/CPU host and instance lifecycle control plane. Bridge firmware, kernel, DPU, and semiconductor architecture with distributed systems. Provide technical leadership, mentor senior engineers, set engineering standards, collaborate with product and data-center teams, and translate customer requirements into scalable infrastructure.
Summary Generated by Built In

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

 

About the Role

As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.
The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.

What You'll Do


We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:

  • Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.

  • Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.

  • Guide design of compute platform multi-tenant security model

  • Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.

  • Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.

  • Work with customers on translating vague customer technical requirements into concrete engineering deliverables.

  • Set engineering standards and lead design reviews for mission-critical cloud software at scale.

Who You are

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.

  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.

  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.

  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.

  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).

  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.

Nice to Have

  • Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).

  • Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)

  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).

  • Experience with Cloud Service Provider Kubernetes offerings.

  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Skills Required

  • 10+ years working on compute control plane distributed systems for deploying and lifecycle managing heterogeneous compute platforms
  • Deep understanding of BIOS/firmware (UEFI), Linux kernel internals, and cradle-to-grave system lifecycle management
  • Deep expertise in durable execution models and distributed systems for cloud-service provisioning
  • Basic knowledge of software-defined networking fundamentals for secure, multi-tenant distributed systems
  • Proven track record of leading large-scale semiconductor hardware enablement and deployment initiatives
  • Proven experience deploying net-new data-centers into a global compute platform
  • Proficiency in one or more programming languages: C/C++, Rust, Python, Go
  • Required presence in Bellevue, San Francisco, or San Jose office 4 days per week
  • Knowledge of Nvidia AI Factory components (GPU hosts, SuperNICs, ConnectX, BlueField DPUs) and related software (DOCA, DOCA SNAP, CUDA)
  • Knowledge of device drivers, virtualization technologies (KVM, QEMU), and kernel-bypass technologies (SR-IOV, DPDK, SPDK)
  • Experience with Cloud Service Provider Kubernetes offerings
  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF)

Lambda Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Lambda and has not been reviewed or approved by Lambda.

  • Fair & Transparent Compensation Pay is considered competitive for an AI infrastructure company, with posted ranges and observed offers indicating strong packages for senior technical roles. Compensation is often characterized as competitive or top‑shelf, aligning with market expectations.
  • Healthcare Strength Health, dental, and vision coverage are characterized as strong, with broad‑network plans noted and positive experiences highlighted. This foundation supports overall satisfaction with core insurance benefits.
  • Leave & Time Off Breadth Flexible or unlimited PTO is described as actually used, complemented by paid holidays and sick time. Generous parental leave examples further expand the time‑off offering.

Lambda Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
106 Employees
Year Founded: 2012

What We Do

Lambda provides computation to accelerate human progress. We're a team of Deep Learning engineers building the world's best GPU workstations and servers. Our products power engineers and researchers at the forefront of human knowledge. Customers include Microsoft, MIT, Los Alamos National Lab, Disney, Tencent, Kaiser Permanente, Stanford, Harvard, Caltech, and the Department of Defense.

Similar Jobs

Lambda Logo Lambda

Staff Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Remote or Hybrid
3 Locations
750 Employees
314K-465K Annually
Remote
United States
501 Employees
212K-286K Annually

Coursera + Udemy  Logo Coursera + Udemy

Principal Product Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
United States
1500 Employees
219K-274K Annually

Mondelēz International Logo Mondelēz International

Product Owner

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
United States
90000 Employees
140K-193K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account