Domain Architect - AI Networking (Nvidia ecosystem)

Posted 12 Hours Ago
Be an Early Applicant
Hiring Remotely in United States
Remote
110K-150K Annually
Expert/Leader
Cloud • Information Technology • Productivity • Security • Software
The Role
Architects and deploys high-performance NVIDIA InfiniBand and Spectrum-X Ethernet fabrics for AI infrastructure. Responsibilities include fabric topology, routing, congestion control, RoCEv2, DPU offload, telemetry, automation, cabling, optical validation, performance testing, and infrastructure-as-code. Leads client delivery, pre-sales scoping, bills of materials, technical workshops, and integration with enterprise networks. The role requires translating complex network architectures into scalable deployments and commercial outcomes.
Summary Generated by Built In
Job Summary & Responsibilities

Minimum Qualifications 

  • 10+ years in networking, HPC, or data center infrastructure engineering, including 7+ years in a customer-facing architecture role. 
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 
  • Deep working knowledge of NVIDIA networking platforms, including Quantum InfiniBand switch architecture (MLNX-OS, UFM) and Spectrum-X Ethernet running NVIDIA Cumulus Linux or SONiC. 
  • Proficiency in RoCEv2 (RDMA over Converged Ethernet), including PFC and ECN, and a working architectural understanding of NVIDIA BlueField DPU and DOCA offload use cases. 
  • Automation and infrastructure-as-code experience, with proficiency in Python and Ansible for switch configuration management. 
  • Demonstrated ability to translate technical architecture into commercial outcomes for senior stakeholders. 

Preferred Qualifications 

  • Industry background: experience within a systems integrator (SI) or managed service provider (MSP) environment. 
  • Multi-vendor exposure: Arista EOS for high-performance AI Ethernet, or Cisco Nexus Dashboard for AI infrastructure. 
  • Routing protocols: solid understanding of BGP, EVPN, and VXLAN for multi-tenant isolation. 
  • Compute integration: understanding of how the network interacts with the host OS (IPoIB, Netlink) and optimization of GPUDirect RDMA (GDR) and NCCL communication patterns. 
  • Scale-across: experience extending an InfiniBand fabric across multiple data center halls (NVIDIA MetroX). 
  • Certifications: NVIDIA-Certified Professional: AI Networking (NCP-AIN); NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO); Arista ACE (Cloud Engineer); Cisco CCNP/CCIE Data Center. 

Certain states and localities require employers to post a reasonable estimate of salary range. A reasonable estimate of the current base pay range for this position is $110,000.00 to $150,000.00 annually. Actual salary will be based on a variety of factors, including shift, location, experience, skill set, performance, licensure and certification, and business needs. The range for this position in other geographic locations may differ. Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that is not included in the base pay.


The well-being of WWT employees is essential. When it comes to our benefits package, WWT has one of the best. We offer the following benefits to all full-time employees:

  • Health and Wellbeing: Health (Medical & Prescription), Dental, and Vision Care, Onsite Health Centers (MO & IL), Employee Assistance Program, Wellness program
  • Financial Benefits: Competitive Pay, Profit Sharing, 401k Plan with Company Matching, Life and Disability Insurance, Flexible Spending Accounts, Tuition Reimbursement
  • Paid Time Off: PTO & Holidays, Parental Leave, Medical Leave, Military Leave, Bereavement, Day of Caring
  • Additional Perks: Family Planning Benefits, Nursing Mothers Benefits, Voluntary Legal, Voluntary Supplemental Accident/Illness/Hospital, Voluntary ID Theft, Pet Insurance, Employee Discount Program

Note: This is not an all-encompassing list and should not be used as a complete description of the plan’s benefits. For more information, see our US benefits website at wwt.com/us-benefits.


We strive to create an environment where all employees are empowered to succeed based on their skills, performance, and dedication. Our goal is to cultivate a culture of belonging that encourages innovation, collaboration, and respect for all team members, ensuring that WWT remains a great place to work for all!


If you require accessibility accommodation(s) or adjustment during any stage of the hiring process, please let your WWT Recruiter know. The recruiter will work with you to understand your needs and help ensure an accessible experience throughout the interview process.


World Wide Technology is an Equal Opportunity Employer.


If you have any questions or concerns about this posting, please email [email protected]


#LI-AF1

#LI-Remote

Preferred Qualifications

World Wide Technology (WWT) strives to make a new world happen. WWT's work benefits clients and partners as much as it does its people and community across the globe.


Founded in 1990, WWT brings together strategy, deep technical expertise and world-class partnerships to help public and private sector organizations design, build and scale intelligent AI, digital, cybersecurity, cloud and infrastructure solutions. Through its Advanced Technology Center (ATC)—a collaborative ecosystem featuring state-of-the-art hardware and software—WWT enables clients and partners to conceptualize, test and validate innovative technology and then deploy solutions at scale using its global integration and distributions capabilities.


With more than 14,000 team members and over 60 locations globally, WWT's culture—grounded in core values and leadership philosophies—has been recognized by Fortune® and Great Place to Work for its commitment to innovation, trust and creating a great place to work for all. WWT provides products and services to large enterprise, global service provider and public sector clients in up to 130 countries across six continents. Softchoice, a World Wide Technology company, supports commercial and SMB markets in the U.S. and Canada.


Want to work with highly motivated individuals on high-performance teams? Join WWT today!


What is the Solutions Consulting & Engineering Team and why join?


Solutions Consulting & Engineering is an organization that is customer-focused and solutions-led. We deliver end-to-end and emerging solutions to drive customer satisfaction and increase profitability and growth. Our world-class management consulting, delivery excellence, and engineering brilliance enable our success. We embody the OneWWT mindset by bringing the right talent at the right time from anywhere within WWT to solve our customer’s problems. Our goal is to bring together business acumen with full-stack technical know-how to develop innovative solutions for our clients’ most complex challenges. 

About the Role 

As Domain Architect - AI Networking, you will be the primary technical authority for the physical and logical lifecycle of high-performance interconnect fabrics across a diverse portfolio of client environments, bridging the gap between architectural design and hands-on execution. You are a builder as much as an advisor: as comfortable configuring adaptive routing on a leaf-spine fabric from the CLI as you are explaining that configuration to a C-level audience. 

As a global systems integrator, we don't simply operate static cloud environments. We design and deliver purpose-built, high-scale AI factories for some of the world's leading enterprises. In this role you will define the reference standard for network infrastructure, moving beyond single-switch administration to architect repeatable, scalable, and automated lossless fabrics. You will act as technical lead on NVIDIA Cloud Partner (NCP) and private enterprise AI cloud deployments, owning the network layer of the compute, network, and storage stack. 

Your time will be split roughly 60/40 between delivering complex AI infrastructure (60%) and providing pre-sales subject matter expertise (40%). You will lead the physical provisioning of InfiniBand and Spectrum-X Ethernet fabrics for NVIDIA DGX SuperPOD, NVIDIA DGX BasePOD, and Cisco AI POD environments, ensuring clients inherit platforms that are genuinely ready for day-2 operations, while helping the sales team scope and cost future deployments. 

Key Responsibilities 

Delivery and implementation 

High-performance fabric design 

  • Architect and deploy non-blocking fat-tree (Clos) topologies using NVIDIA Quantum-2 (NDR) and Quantum-3 (XDR) InfiniBand switches, and NVIDIA Spectrum-4 Ethernet switches. 
  • Implement rail-optimized network designs so GPU-to-GPU traffic aligns with compute-node PCIe topology, minimizing latency for NCCL collective operations. 
  • Configure adaptive routing, congestion control, and quality of service (QoS) to prevent head-of-line blocking and guarantee lossless data delivery. 

In-network computing and offload strategy 

  • InfiniBand: enable and tune SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) on Quantum switches to offload collective operations (AllReduce, ReduceScatter) from the GPU to the switch silicon. 
  • Ethernet: architect high-performance Ethernet AI fabrics using SmartNIC/DPU offload (NVIDIA BlueField-3) to accelerate collective operations and isolate management traffic from the data path. 
  • Evaluate and roadmap emerging Ultra Ethernet Consortium (UEC) standards as the practice transitions in-network collectives from proprietary InfiniBand toward open Ethernet. 
  • Tune RoCEv2, PFC, and ECN on Ethernet fabrics and SmartNICs to approximate InfiniBand-class lossless behavior. 

Fabric management, automation, and telemetry 

  • Deploy NVIDIA UFM (Unified Fabric Manager) to manage InfiniBand subnets and NetQ for Ethernet fabric telemetry; implement PTP for nanosecond-level clock synchronization across the cluster. 
  • Use automation and AI-assisted tooling to generate switch configuration from the P2P cabling schedule, LLD, and PDG, then validate every topology in NVIDIA DSX Air before it ever touches hardware. 
  • Monitor fabric health continuously to identify slow receivers and link degradation in real time, and maintain Python and Ansible-based infrastructure-as-code for ongoing configuration management. 

Layer 1 precision and performance engineering 

  • Own the physical cabling strategy, defining cable schedules (DAC, AOC, or OSFP transceivers) that meet signal-integrity requirements over the required distances. 
  • Validate the link budget to ensure optical loss stays within acceptable limits for 400G and 800G links. 
  • Conduct acceptance and validation testing using NCCL tests to verify fabric performance against expected baselines. 

Pre-sales SME and consulting 

Technical scoping and estimation 

  • Calculate the bisectional bandwidth required for a client's workload (for example, training typically requires 1:1 non-blocking; inference may tolerate 3:1 oversubscription), and produce accurate level-of-effort (LOE) estimates for statements of work. 

Bill of materials validation and architecture 

  • Own the technical accuracy of the network bill of materials, validating that every switch, transceiver, and cable is on the NVIDIA or OEM hardware compatibility list (HCL). 
  • Manage the complexity of breakout cables (for example, 800G to 2x400G) and connector types (OSFP, QSFP112) to prevent on-site installation failures. 

Client workshops 

  • Educate clients on the difference between standard enterprise Ethernet (lossy, TCP-based) and AI fabric (lossless, RDMA-based) traffic. 
  • Design the integration point between the high-speed “back-end” AI fabric and the client's existing “front-end” management network, including BGP/EVPN handoffs. 

Skills Required

  • 10+ years of experience in networking, HPC, or data center infrastructure engineering
  • 7+ years of experience in a customer-facing architecture role
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • Deep knowledge of NVIDIA networking platforms, including Quantum InfiniBand, MLNX-OS, UFM, Spectrum-X Ethernet, Cumulus Linux, or SONiC
  • Proficiency in RoCEv2, PFC, ECN, NVIDIA BlueField DPU, and DOCA offload use cases
  • Automation and infrastructure-as-code experience with Python and Ansible
  • Ability to translate technical architecture into commercial outcomes for senior stakeholders
  • Experience in a systems integrator or managed service provider environment
  • Experience with Arista EOS or Cisco Nexus Dashboard for AI infrastructure
  • Understanding of BGP, EVPN, and VXLAN
  • Understanding of IPoIB, Netlink, GPUDirect RDMA, and NCCL communication patterns
  • Experience extending InfiniBand fabrics across multiple data center halls using NVIDIA MetroX
  • NVIDIA NCP-AIN, NCA-AIIO, Arista ACE, Cisco CCNP Data Center, or Cisco CCIE Data Center certification
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chicago, IL

Similar Jobs

World Wide Technology Logo World Wide Technology

Architect

Big Data • Cloud • Hardware • Software • App development
Remote
United States
9000 Employees
110K-150K Annually

Zscaler Logo Zscaler

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
Florida, USA
8697 Employees
120K-170K Annually
Remote
USA
62 Employees
90K-110K Annually

ServiceNow Logo ServiceNow

Enterprise Account Exec (Armis/Veza)

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Denver, CO, USA
29000 Employees
114K-165K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account