Staff Product Manager - Compute

Posted Yesterday
Be an Early Applicant
2 Locations
Remote or Hybrid
291K-430K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Own the compute product strategy for GPU and CPU infrastructure, including instance offerings, capacity, provisioning, placement, reliability, performance isolation, pricing, and commitments. Partner with engineering and cross-functional teams to define requirements, prioritize work, launch infrastructure capabilities, and measure customer adoption and utilization. The role requires deep cloud compute experience, customer insight, commercial judgment, and ownership of operational behavior after launch.
Summary Generated by Built In

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

Note: This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda's designated work from home day is currently Tuesday.

About the Role

Compute is the product at Lambda. Customers come to us for graphics processing units (GPUs) they can get, hold, and run hard, and nearly everything else we sell depends on the compute layer working well.

As a product manager on the Compute team, you will own a meaningful part of defining the future of how we provide compute to our customers. The domain covers the shape of what customers rent, meaning instance families, sizes, and tenancy, across both GPUs and central processing units (CPUs). It covers how customers get capacity and hold onto it. It covers what they build production automation against, and how fast and predictably that behaves. It covers what happens once the workload is running, which means placement, performance isolation, hardware failure, maintenance, and the signals customers need to run their own operations. You will work across these areas depending on which customer needs are burning hottest.

Most of this is the feature set a mature compute cloud already has and Lambda does not have yet. A neocloud is not a hyperscaler, though, and copying that list wholesale is the wrong instinct. Part of the job is judgment about which of those capabilities matter here and which carry cost we should not pay. The other part is deciding where serving AI workloads well means building something the hyperscalers never needed.

Great product managers at Lambda are defined by three things: insight, influence, and execution. Insight means you look at the data, determine what it means for customers and business, and then figure out what to do about it. But, a great idea doesn't mean anything in a vacuum. That is where influence comes in. Influence means you take that idea and get others to want to buy into it; you win over engineers, designers, executives, and partners without relying on authority. But a great idea that everyone is excited about doesn't matter unless it is delivered to customers. Execution means you work with the right people to get the idea launched, then measure and iterate. We hire product managers who learn new domains fast and reason rigorously from evidence. Deep compute platform experience at a hyperscaler or neocloud, across both GPUs and CPUs, is highly desired.

If you have built compute primitives at a cloud provider and want to do it again somewhere the answers are not settled, we'd love to hear from you.

We value diverse backgrounds, experiences, and skills, and we are excited to hear from candidates who can bring unique perspectives to our team. If you do not exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role.

What You'll Do

  • Find the Real Problem: Get close to customers, deals, escalations, and support on On-Demand GPU Instances and 1-Click Clusters, and work out what is actually blocking them rather than what they asked for.

  • Decide What Matters: Rank the work and sequence it, and be honest about what Lambda is not doing this year and why.

  • Write the Definition: Produce requirements, user stories, and acceptance criteria precise enough that engineering builds from them without a translation layer.

  • Price and Package It: Decide how your part of compute is priced, packaged, and committed to across NVIDIA GPU generations, and how customers compare the options.

  • Deliver With Engineering: Work through the build with the compute and control plane teams, make the tradeoff calls that come up mid-flight, and keep scope honest against the date.

  • Land the Launch: Set launch criteria that cover the operational readiness an infrastructure product needs, and get documentation, pricing, sales, and support in place before it goes live rather than after.

  • Measure and Iterate: Define what success looks like before launch, then go find out whether it happened, using adoption and utilization rather than opinion.

  • Make the Call Stick: Take a position on contested tradeoffs, write it down well enough that people can disagree with it precisely, and keep owning the decision after it is made.

  • Work as a Group: Partner with the other product managers on compute and with the teams that own storage, networking, orchestration, and commerce, so the pieces land as one product.

You

  • Have 7+ years of product management experience, including time on a compute platform at a hyperscaler, neocloud, or comparable cloud provider.

  • Have shipped compute primitives that external customers used at scale, like instance types, capacity products, provisioning interfaces, or placement and isolation controls.

  • Understand GPU and CPU infrastructure well enough to reason with engineers about launch paths, hardware failure modes, placement, and performance.

  • Have owned a product where day two behavior mattered as much as launch, including failure handling, maintenance, and what customers are told when something breaks.

  • Have owned pricing, packaging, or commitment terms for an infrastructure product.

  • Can tell which parts of a hyperscaler playbook transfer to an AI cloud and which are cost without benefit.

  • Can reason about utilization and capacity economics, and make the call when customer flexibility and hardware utilization pull against each other.

  • Able to define iterative plans that move an organization from the current state towards the desired outcome.

  • Can turn ambiguous customer, technical, and commercial inputs into product definition that multiple teams execute against.

  • Communicate plainly, write well, and make decisions easier for people who do not all share the same context.

Nice to Have

  • Experience with interruptible or preemptible capacity, or with reservations and committed use products.

  • Experience with the surfaces customers automate against, like public APIs, infrastructure as code providers, machine images, or snapshots.

  • Experience with hardware failure handling across a large hardware footprint, like spare pools, node replacement, or fault containment.

  • Experience with performance isolation and placement on shared hardware, including what a provider can honestly guarantee.

  • Experience with bare metal or dedicated host products, including isolation between tenants.

  • Experience with distributed training or large-scale inference workloads and what they demand of the compute layer.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Skills Required

  • 7+ years of product management experience, including experience with a compute platform at a hyperscaler, neocloud, or comparable cloud provider
  • Experience shipping compute primitives used at scale by external customers, such as instance types, capacity products, provisioning interfaces, or placement and isolation controls
  • Strong understanding of GPU and CPU infrastructure, including launch paths, hardware failure modes, placement, and performance
  • Experience owning products where failure handling, maintenance, and customer communication after launch were important
  • Experience owning pricing, packaging, or commitment terms for an infrastructure product
  • Ability to evaluate which hyperscaler capabilities transfer effectively to an AI cloud
  • Ability to reason about utilization and capacity economics and balance customer flexibility against hardware utilization
  • Ability to define iterative plans from the current state toward a desired outcome
  • Ability to translate ambiguous customer, technical, and commercial inputs into product definitions for multiple teams
  • Clear, plain communication and strong writing skills
  • Experience with interruptible or preemptible capacity, reservations, or committed-use products
  • Experience with customer automation surfaces such as public APIs, infrastructure-as-code providers, machine images, or snapshots
  • Experience with hardware failure handling across large hardware footprints
  • Experience with performance isolation and placement on shared hardware
  • Experience with bare-metal or dedicated-host products and tenant isolation
  • Experience with distributed training or large-scale inference workloads

Lambda Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Lambda and has not been reviewed or approved by Lambda.

  • Fair & Transparent Compensation — Compensation is described as competitive or generous, particularly in technical roles. Some employees also highlight strong cash pay relative to workload and remote flexibility.
  • Healthcare Strength — Healthcare coverage is described as comprehensive, with medical, dental, and vision benefits that extend to dependents. This core coverage is positioned as a strength alongside competitive pay.
  • Leave & Time Off Breadth — Time off policies include flexible PTO that is encouraged and used, plus roughly 12 paid holidays and around 5 sick days. These elements are noted as helping reduce burnout and supporting flexibility.

Lambda Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
750 Employees
Year Founded: 2012

What We Do

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure and the #1 GPU Cloud for ML/AI teams. Their mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

Similar Jobs

Remote or Hybrid
2 Locations
106 Employees
291K-430K Annually

Compa Logo Compa

Vice President Of Finance

Artificial Intelligence • HR Tech • Software • Business Intelligence
Remote or Hybrid
Office, Machaze, Manica, MOZ
75 Employees
200K-250K Annually

Mondelēz International Logo Mondelēz International

Senior Product Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
3 Locations
90000 Employees

Mondelēz International Logo Mondelēz International

o9 Data & Integration Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
2 Locations
90000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account