Network Infrastructure Architect

Reposted 2 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
250K-330K Annually
Expert/Leader
Artificial Intelligence • Hardware • Machine Learning • Natural Language Processing • Software • Generative AI
SambaNova is the #1 platform for business AI.
The Role
Design and define hyperscale AI compute network architectures. Translate customer and workload needs into scalable, cost‑optimized network and compute product requirements. Lead cross-functional integration, fabric design (InfiniBand/RoCE/Ethernet), lossless Ethernet tuning, high‑speed optics/cabling decisions, lifecycle operations, and deployment strategies for large AI clusters.
Summary Generated by Built In

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About the role

SambaNova is accelerating from rack-scale to global-scale AI compute clusters. We're looking for a Network Architect join us on this exciting new chapter and own this critical layer end to end.

Responsibilities

In this role, you'll be architecting the network fabrics that power hyperscale AI clusters, from RDU-accelerated servers to global-scale deployments, owning the fabric, interconnect, and spine-leaf design decisions that shape how our compute platforms are built. You'll partner closely with hardware engineering, supply chain, and customer-facing teams to translate workload requirements into platform roadmaps, taking a vendor-agnostic approach that avoids lock-in. Your architecture decisions will directly shape the products we ship and how fast the world's largest AI clusters come online.

Required Qualifications
  • 12+ years designing or architecting infrastructure for hyperscale, AI, HPC, or large-scale data center environments
  • Deep working knowledge of AI cluster networking, including InfiniBand, RoCEv2, Ethernet-based AI fabrics, 400G/800G interconnects, and GPU-to-GPU east-west traffic patterns
  • Experience with lossless Ethernet designs, including QoS, ECN, PFC, congestion management, buffer tuning, and failure-domain isolation
  • Strong understanding of spine-leaf architectures and associated technologies (VXLAN, EVPN, BGP, ECMP, underlay/overlay design, network automation)
  • Familiarity with high-speed optics and cabling for AI clusters, including 400G/800G transceivers, DAC/AOC, fiber topology, and link budgets
  • Experience translating customer and workload requirements into product requirements, platform roadmaps, and architecture tradeoffs
  • Vendor-agnostic approach to platform selection and interoperability across switch, NIC, optical, and fabric ecosystems
  • Ability to partner across hardware engineering, supply chain, operations, and customer-facing teams to bring platforms from concept through deployment
Preferred Qualifications
  • Experience with AI compute platforms: GPU/accelerator systems, rack-scale compute, NVLink/NVSwitch-class fabrics, PCIe/CXL, DPUs/IPUs
  • Understanding of rack-level power, thermal, mechanical, and serviceability requirements, and how network choices drive them (NIC selection, cabling density, PCIe lane allocation)
  • Background in large-scale fleet operations: lifecycle management, reliability, observability, telemetry, field issue resolution

Base Salary Range:

Base Pay Range
$250,000$330,000 USD

Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Skills Required

  • 12+ years designing, architecting, or productizing compute infrastructure for hyperscale, AI, HPC, cloud, or large-scale data center environments
  • Deep experience with AI compute platforms including GPU/accelerator systems, high-density server architectures, rack-scale compute, and cluster-level design
  • Strong understanding of server and rack-level architecture: power, thermal, mechanical, serviceability, firmware, and platform integration
  • Translate customer, workload, and deployment requirements into compute product requirements, platform roadmaps, technical specifications, and architecture tradeoffs
  • Familiarity with modern AI infrastructure components: GPU servers, accelerator trays, NVLink/NVSwitch-class fabrics, PCIe/CXL, high-speed NICs, DPUs/IPUs, storage-attached compute
  • Deep working knowledge of AI cluster networking and fabric dependencies including InfiniBand, RoCEv2, Ethernet-based AI fabrics, and 400G/800G interconnects
  • Experience with lossless or near-lossless Ethernet designs (QoS, priority mapping, ECN, PFC, congestion management, buffer tuning, telemetry, failure-domain isolation)
  • Strong understanding of spine-leaf architectures and related technologies (VXLAN, EVPN, BGP, MLAG, ECMP, underlay/overlay, multi-tenant segmentation, network automation)
  • Vendor-agnostic approach to architecture, platform selection, interoperability, and lifecycle strategy
  • Experience comparing and integrating technologies across switches, NICs, accelerators, optics, and fabric ecosystems to avoid vendor lock-in
  • Familiarity with high-speed optics and cabling for AI clusters (400G/800G DR4/FR4/LR4, SR, DAC/AOC, fiber topology, link budgets, transceiver interoperability)
  • Ability to evaluate network choices' impact on compute product requirements (NIC selection, DPU/IPU integration, PCIe lanes, rack power/thermal, cabling density, latency, throughput, resiliency, serviceability)
  • Partner closely with hardware engineering, supply chain, manufacturing, operations, networking, facilities, vendors/OEMs/ODMs, and customer-facing teams from concept through deployment
  • Background supporting large-scale infrastructure operations: fleet deployment, lifecycle management, reliability, observability, telemetry, automation, and field issue resolution
  • Strong product mindset balancing performance, cost, manufacturability, deployment velocity, serviceability, supply availability, power efficiency, ecosystem flexibility, and scalability
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
500 Employees
Year Founded: 2017

What We Do

AI is changing the world and at SambaNova, we believe that you don’t need unlimited resources to take advantage of the most advanced, valuable AI capabilities - capabilities that are helping organizations explore the universe, find cures for cancer, and giving companies access to insights that provide a competitive edge. We deliver the world’s fastest and only complete AI solution for enterprises and governments with world-record inference performance and accuracy. Powered by the SambaNova SN40L Reconfigurable Dataflow Unit (RDU), organizations can build a technology backbone for the next decade of AI innovation with SambaNova Suite. Our fully integrated hardware-software system, DataScale®, enables organizations to train, fine-tune, and deploy the most demanding AI workloads using the largest and most challenging models. Most recently, with the launch of our newest offering, SambaNova Cloud, developers can supercharge AI-powered applications on Llama 3.2 models. SambaNova was founded in 2017 in Palo Alto, California, by a group of industry luminaries, business leaders, and world-class innovators who understand AI. Today, we’ve built an incredibly smart and motivated team dedicated to making a lasting impact on the industry and equipping our customers to thrive in the new era of AI.

Why Work With Us

As a talent first company, we aim to hire the greatest and most innovative minds in the industry- driving the next generation of AI computing where no barrier is too high and the possibilities are truly limitless. We encourage our peers to take risks and take the initiative to make a lasting impact on the AI and ML industries.

Gallery

Gallery

Similar Jobs

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Lead Supervisor I

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Hillgrove, CA, USA
16000 Employees
17-28 Hourly

Collectors Logo Collectors

Demand Planner

Consumer Web • eCommerce • Machine Learning • Software • Sports • Analytics
In-Office
Santa Ana, CA, USA
2246 Employees
107K-144K Annually

Arm Logo Arm

Senior Manager Revenue Forecasting

Artificial Intelligence • Internet of Things • Semiconductor
Hybrid
San Jose, CA, USA
8314 Employees
187K-253K Annually

Metropolis Technologies Logo Metropolis Technologies

Associate Account Executive

Artificial Intelligence • Computer Vision • Machine Learning • Payments • Real Estate • PropTech
Easy Apply
In-Office
San Francisco, CA, USA
23100 Employees
85K-100K Annually

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account