Director, Network Capacity Automation

Posted 22 Hours Ago
Be an Early Applicant
2 Locations
Remote or Hybrid
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Leads network capacity automation for Lambda’s AI cloud infrastructure. Owns capacity planning, delivery, lifecycle, forecasting, and operational reliability while building software to replace manual processes. Hires and develops engineering leaders and teams, sets multi-year technical strategy, partners across infrastructure, supply chain, data center operations, product, finance, and interconnect teams, and represents technical direction to senior leadership and customers.
Summary Generated by Built In

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

Our vision is bold and is not an incremental exercise. We will continually re-evaluate and reinvent our current ways of working and the automation underneath all of it, while operating the existing network flawlessly for all customer workloads.

Lambda has 10x'd over the last three years, and the network engineering organization is scaling to match. We're looking for a Director of Network Capacity Automation to lead one or more teams building the software that enables us to plan and deliver network capacity ahead of customer demand. This capacity is critical to ensure a flawless customer connectivity experience to our GPU cloud infrastructure. We're looking for a Director to own that outcome end to end — and to build the organization that delivers it. You'll hire and grow a talented team of software and network proficient engineers, shape team culture, and partner with technical leadership to deliver the mechanisms needed to stay ahead of our scale ambitions.

This role will report to the VP of Cloud and AI Networking, working alongside a highly capable and passionate team of engineers.

If you'd like to join a great team of customer obsessed engineers building the world's best AI cloud, come join us.


What You'll Do

  • Own network capacity end to end — planning, delivery, turn-up and lifecycle — and be accountable for capacity landing ahead of demand

  • Build and lead the organization that delivers it: hire, develop and retain engineers and the leaders who manage them, and design the team structure as the org grows

  • Set the multi-year technical strategy for capacity automation to ensure it continually scales

  • Convert manual process into durable software — the goal is a capacity pipeline that runs without heroics, not a better-organised set of runbooks

  • Own the operational bar for your org: how capacity work is planned, reviewed, shipped, measured and learned from when it goes wrong

  • Manage through leads and senior ICs, developing both — regular 1:1s, clear performance feedback, growth planning, and sponsorship of meaningful work

  • Partner and support our Interconnect team who negotiate, acquire and coordinate on external connectivity

  • Partner with Principal Engineers and technical leadership to maintain a high engineering bar across design, code quality and operational reliability

  • Partner with Infrastructure, Supply Chain, Data Center Operations, Product and Finance to align capacity roadmaps, forecast demand and manage long-lead dependencies

  • Represent your org's work and technical direction to senior leadership, and to enterprise customers and partners where appropriate

  • Contribute to org-wide engineering process improvements — how we plan, how we ship, how we learn from incidents

You

  • Have 8+ years of engineering management experience, including directly managing senior individual contributors

  • Proven track record of building and scaling engineering organizations that deliver mission-critical, high-performance cloud connectivity serving millions of users

  • Have a software engineering background — you've shipped production systems and can engage credibly on architecture, technical tradeoffs, and code quality

  • Foster a data driven, automation first, high velocity organization

  • Familiarity designing, building, scaling and operating large-scale networks

  • Familiarity with how to build distributed systems and large-scale networking services

  • Strong automation skills and track record accelerating a business delivery velocity on a large scale

  • Have a track record of building high-performing teams in fast-moving, technically demanding environments

  • Are skilled at translating ambiguous business and product goals into clear team priorities and executable engineering plans

  • Show strong judgment about when to go deep technically, when to delegate, and when to escalate

  • Track record of successful project and product delivery in fast-paced, high-pressure environments

  • Excellent communicator both written and verbal. You leverage high quality written artifacts to inform high quality decision making and get approval on problems and solution opportunities

Nice to Have

  • Experience in GPU cloud, HPC, or AI/ML infrastructure environments

  • Deep familiarity with large-scale data center or cloud networking — fabric build-out, capacity modelling, or network supply chain

  • Experience with infrastructure capacity forecasting and long-lead hardware planning

  • Prior experience at a high-growth infrastructure or cloud company

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Skills Required

  • 8+ years of engineering management experience
  • Experience directly managing senior individual contributors
  • Track record building and scaling engineering organizations delivering mission-critical, high-performance cloud connectivity serving millions of users
  • Software engineering background with production systems experience
  • Ability to engage credibly on architecture, technical tradeoffs, and code quality
  • Experience designing, building, scaling, and operating large-scale networks
  • Experience building distributed systems and large-scale networking services
  • Strong automation skills and a record of accelerating business delivery velocity at scale
  • Track record building high-performing teams in fast-moving, technically demanding environments
  • Ability to translate ambiguous business and product goals into clear priorities and executable engineering plans
  • Strong technical judgment regarding when to go deep, delegate, or escalate
  • Track record of successful project and product delivery in fast-paced, high-pressure environments
  • Excellent written and verbal communication skills
  • Experience in GPU cloud, HPC, or AI/ML infrastructure environments
  • Deep familiarity with large-scale data center or cloud networking, including fabric build-out, capacity modeling, or network supply chain
  • Experience with infrastructure capacity forecasting and long-lead hardware planning
  • Prior experience at a high-growth infrastructure or cloud company
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
750 Employees
Year Founded: 2012

What We Do

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure and the #1 GPU Cloud for ML/AI teams. Their mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

Similar Jobs

Remote or Hybrid
2 Locations
106 Employees

Mondelēz International Logo Mondelēz International

Scientist

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
3 Locations
90000 Employees

Mondelēz International Logo Mondelēz International

R&D Manager - Cote d'Or Chocolate

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
4 Locations
90000 Employees

Invenergy Logo Invenergy

Senior PI Administrator

Greentech • Real Estate • Social Impact • Energy • Industrial • Solar • Renewable Energy
Remote or Hybrid
18 Locations
2500 Employees
125K-155K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account