Senior Staff Network Engineer

Posted 3 Days Ago
2 Locations
Remote or Hybrid
324K-480K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Lead architecture, design, scaling, automation, and operation of Lambda’s large-scale AI cloud network. Own major network domains, lead cross-team projects, resolve complex production issues, qualify hardware and software, improve observability and self-healing, mentor engineers, shape technical strategy, and participate in on-call operations. The role requires deep expertise in data center, backbone, internet, cloud, Linux, network automation, and multi-vendor networking technologies.
Summary Generated by Built In

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

You will be joining a team of software, hardware and network engineers building one of the largest AI training and inference networks in the world. Lambda has 10x'd over the last three years and the network has to keep pace.

As a Senior Staff Network Engineer you'll have the opportunity to be across all aspects of the network, including industry leading backend, frontend, backbone and edge networks. You will lead the projects that design, build and scale it. You are the engineer the team routes its hardest problems to, and the one whose designs other engineers follow.

Our vision is bold and is not an incremental exercise. We will continually re-evaluate and reinvent our current fabric topology, routing design and the automation underneath all of it, while operating the existing network flawlessly for all customer workloads.

This role will report to our Senior Network Engineering leader and will work alongside an existing, highly capable team of network engineers.


What You’ll Do

  • Own the architecture of major areas of Lambda's network end to end — you make the design calls, you help build it, and you're connected to how they hold up in production

  • Lead large networking projects that may span multiple teams, from design through build to turn-up, and drive them to completion

  • Write the designs, RFCs and reviews other engineers build from, and hold the bar in design review for work across the team

  • Be the deep technical escalation point for the hardest production problems — the packet path, the routing table, the vendor's firmware, the code

  • Drive the automation of your domain, replacing manual operational work with software that holds up under scale

  • Partner with hardware, platform and product teams to resolve technical questions that cross team boundaries

  • Qualify and integrate new network hardware and software, and hold vendors to account on defects and roadmap

  • Guide partner teams to automate and scale through robust software systems and services

  • Guide partner teams build the right telemetry and automation to ensure our network health is observable and self healing

  • Mentor senior and mid-level engineers, and raise the bar through design review, code review and hiring

  • Contribute to the multi-year technical strategy for Lambda's network, and own the parts of it that fall in your area

  • Work with internal and external customers to resolve network related issues

  • Operate what you build. You'll take part in day 2 operations and the on-call rotation, because the fastest feedback loop between design and reality runs through the pager

You

  • Have owned the technical direction of a significant area of a production network — you can walk us through a design you made, why, what it cost, and what you'd change

  • Have led large production-scale networking projects across multiple teams, from design through delivery

  • Expert in large network designs to achieve maximum availability and performance

  • Deep expertise in data center, backbone and internet protocols and technology

  • Experience building and tuning networks for low latency, fast convergence and high availability

  • Have experience with cloud provider networking (such as AWS, GCP, OCI)

  • Are comfortable on the Linux command line, and have an understanding of the Linux networking stack and internals

  • Strong automation skills (Python, Ansible, Jinja), network APIs, and git or similar source control

  • Production experience with multiple network gear vendors (Arista, Juniper, Cisco, Cumulus/SONiC, Opengear)

  • Have raised the technical bar of engineers around you through mentorship and design review

  • Excellent communicator both written and verbal. You leverage high quality written artifacts to inform high quality decision making and get approval on problems and solution opportunities

Nice to Have

  • Have knowledge or experience maintaining Software Defined Networks (SDN)

  • Experience automating network configuration

  • Hands-on with HPC/AI networking: RoCEv2 and/or InfiniBand (Congestion Control, VLs, partitions), GPUDirect RDMA concepts.

  • Experience with DWDM technologies and SD-WAN

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Skills Required

  • Significant experience owning the technical direction of a production network area
  • Experience leading large production-scale networking projects across multiple teams from design through delivery
  • Expertise in large network designs optimized for availability and performance
  • Deep expertise in data center, backbone, and internet protocols and technologies
  • Experience building and tuning networks for low latency, fast convergence, and high availability
  • Experience with cloud provider networking, such as AWS, GCP, or OCI
  • Comfort with the Linux command line and understanding of the Linux networking stack and internals
  • Strong automation skills using Python, Ansible, and Jinja
  • Experience with network APIs and Git or similar source control
  • Production experience with multiple network gear vendors, including Arista, Juniper, Cisco, Cumulus or SONiC, and Opengear
  • Experience mentoring engineers and conducting design reviews
  • Excellent written and verbal communication skills
  • Knowledge or experience maintaining Software Defined Networks
  • Experience automating network configuration
  • Hands-on experience with HPC or AI networking, including RoCEv2, InfiniBand, congestion control, VLs, partitions, or GPUDirect RDMA
  • Experience with DWDM technologies and SD-WAN
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
750 Employees
Year Founded: 2012

What We Do

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure and the #1 GPU Cloud for ML/AI teams. Their mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

Similar Jobs

Remote or Hybrid
2 Locations
106 Employees
324K-480K Annually

Mondelēz International Logo Mondelēz International

o9 Data & Integration Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
2 Locations
90000 Employees

Compa Logo Compa

Enterprise Account Executive

Artificial Intelligence • HR Tech • Software • Business Intelligence
Remote or Hybrid
3 Locations
75 Employees
200K-225K Annually

Mondelēz International Logo Mondelēz International

Scientist

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
3 Locations
90000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account