Network Automation Software Lead

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Las Vegas, NV, USA
In-Office or Remote
Senior level
Artificial Intelligence • Cloud • Software
The Role
Lead development of a zero-touch provisioning and network automation platform for GPU datacenter fabrics. Own ZTP pipeline, intent-based config, validation/digital-twin testing, telemetry/observability, APIs, and team hiring/mentorship while contributing code, CI/CD, and on-call ownership.
Summary Generated by Built In

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

 

About the Role

We're hiring a Network Automation Software Lead to build and own the end-to-end zero-touch provisioning (ZTP) and automation platform that stands up and operates our GPU network fabrics and to lead the small team of engineers and SREs building it.

The goal is zero human intervention and full fabric validation: a switch goes from rack-and-stack to production-ready automatically, ensuring every link across the fabric is validated against the plan and spec. Our Network Engineering team is your customer; you build the platform, they run the network on top of it.

You'll report to a Software Engineering Lead, which means this is real software engineering, not scripts bolted onto a NOC. Source control, testing, CI/CD, code review, release management, and on-call are the baseline. You'll also be responsible for making sure the platform integrates cleanly with the rest of our software and platform tooling.

 

What You’ll Do

  • Own the end-to-end ZTP pipeline: bare-metal switch boot → image + base config → registration in source of truth → full intended config → validation → production — with zero human intervention.

  • Build intent-based config generation off a network source of truth / IPAM, with GitOps-style deployment, pre/post-change validation, and safe rollout and rollback.

  • Establish network validation and pre-deployment testing (snapshot/digital-twin testing) so changes are caught before they hit production fabrics.

  • Build streaming telemetry and metrics/logging pipelines (gNMI / OpenConfig) for fabric health.

  • Instrument what matters for GPU networks: RoCE health (PFC/ECN counters), optics and link errors, BGP / EVPN state, capacity and utilization.

  • Deliver dashboards and alerting the network team actually uses — signal, not noise.

  • Gather requirements, build self-service APIs and interfaces, and relentlessly accelerate their deployment velocity.

  • Partner closely so the tooling reflects how the network is actually operated and turned up.

  • Hire, mentor, and grow a small team of software engineers and SREs; own roadmap, prioritization, and delivery — while still carrying a meaningful share of the code yourself.

  • Set technical direction and standards, and ensure clean integration points with the broader platform stack (infra provisioning, CI/CD, secrets, identity, existing observability).

  • Bring software engineering rigor to network automation: code review, testing, release management, and on-call ownership.

 

Who You Are

Required Qualifications

  • 8+ years of relevant experience

  • Proven experience building network automation at scale — ideally at a hyperscaler, large cloud, or large-scale datacenter / AI-infrastructure operator.

  • You've built or been a core contributor to a ZTP / device-provisioning system end-to-end, not just maintained one.

  • Strong software engineering fundamentals: Python and/or Go, with real production practices (version control, testing, CI/CD, code review).

  • Hands-on depth with datacenter Clos fabrics and the protocols that run them: BGP, EVPN/VXLAN, and ideally RoCEv2 / RDMA for GPU networks at scale.

  • Fluency with modern network automation tech: gNMI/gNOI, OpenConfig/YANG, NETCONF; source-of-truth systems (NetBox / Nautobot); NOS platforms (SONiC/FRR or vendor equivalents); tooling like Nornir / NAPALM / Ansible.

  • Experience with observability / telemetry pipelines (Prometheus, Grafana, Kafka, OpenTelemetry, or similar).

  • Comfort running services on Kubernetes / containers.

  • Leadership: you've led a team or been the clear technical owner of a platform, and you instinctively treat internal users as customers.

Preferred Qualifications

  • Experience with GPU / AI training or inference clusters and their backend networks.

  • Familiarity with the AMD networking ecosystem (Pensando DPUs, Ultra Ethernet) or building on Ethernet-based RDMA fabrics.

  • Whitebox / disaggregated networking and SONiC at scale.

  • Network validation / digital-twin tooling (e.g., Batfish, containerlab).

  • Multi-site / multi-region datacenter buildouts.

 

What We Offer

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

 

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

 

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].

 

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in United States, as required by law.

 

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

 

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Skills Required

  • 8+ years of relevant experience
  • Proven experience building network automation at scale (hyperscaler/datacenter/operator)
  • Built or been a core contributor to a ZTP / device-provisioning system end-to-end
  • Strong software engineering fundamentals and production practices (version control, testing, CI/CD, code review)
  • Production experience in Python and/or Go
  • Hands-on depth with datacenter Clos fabrics and protocols: BGP, EVPN/VXLAN, and RoCEv2/RDMA
  • Fluency with network automation tech: gNMI, gNOI, OpenConfig, YANG, NETCONF
  • Familiarity with source-of-truth systems / IPAM (NetBox, Nautobot)
  • Experience with NOS platforms (SONiC, FRR or vendor equivalents)
  • Experience with automation tooling (Nornir, NAPALM, Ansible)
  • Experience with observability/telemetry pipelines (Prometheus, Grafana, Kafka, OpenTelemetry or similar)
  • Comfort running services on Kubernetes / containers
  • Leadership experience: led a team or been the technical owner of a platform
  • Experience with GPU / AI training or inference clusters and their backend networks
  • Familiarity with AMD networking ecosystem (Pensando DPUs, Ultra Ethernet) or Ethernet-based RDMA fabrics
  • Whitebox / disaggregated networking and SONiC at scale
  • Network validation / digital-twin tooling (Batfish, containerlab)
  • Experience with multi-site / multi-region datacenter buildouts
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Las Vegas, Nevada
56 Employees

What We Do

TensorWave is a cutting-edge cloud platform designed specifically for AI workloads. Offering AMD MI300X accelerators and a best-in-class inference engine, TensorWave is a top-choice for training, fine-tuning, and inference. Visit tensorwave.com to learn more. Send us a message to try it for free.

Similar Jobs

Headway Logo Headway

Senior Manager, Content & Community Marketing

Consumer Web • Healthtech • Professional Services • Social Impact • Software
Remote
USA
819 Employees
180K-225K Annually

Coinbase Logo Coinbase

Senior Recruiter

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
USA
4700 Employees
131K-154K Annually
Remote
United States
40 Employees
75K-125K Annually
Easy Apply
Remote
United States
900 Employees
110K-127K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account