Domain Architect - AI Compute (Nvidia Ecosystem)

Posted 13 Hours Ago
Be an Early Applicant
Hiring Remotely in United States
Remote
110K-150K Annually
Expert/Leader
Big Data • Cloud • Hardware • Software • App development
WWT makes a new world happen.
The Role
Architect and deliver high-scale NVIDIA GPU compute environments for enterprise AI workloads. Lead bare-metal provisioning, cluster management, operating-system hardening, scheduler and workload configuration, monitoring, performance validation, and day-two operations. Serve as a customer-facing technical authority, supporting presales discovery, architecture, bills of materials, estimation, and client workshops. Design repeatable AI infrastructure across DGX, HGX, MGX, NVL72, and Cisco AI POD platforms.
Summary Generated by Built In
Job Summary & Responsibilities

Minimum Qualifications 

  • 10+ years in HPC, AI, or data center infrastructure engineering, including 7+ years in a customer-facing architecture role. 
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 
  • Deep working knowledge of NVIDIA GPU platforms across the Hopper and Blackwell architectures, including Grace Hopper and Grace Blackwell superchip systems, and the associated software stack (CUDA, cuDNN, NCCL). 
  • Strong Linux systems engineering skills, including distributions tuned for HPC and AI workloads (Ubuntu, RHEL), kernel tuning, driver management, and system hardening. 
  • Automation and infrastructure-as-code experience, with proficiency in Ansible and Python, and familiarity with declarative tooling such as Terraform. 
  • Demonstrated ability to translate technical architecture into commercial outcomes for senior stakeholders. 

 

Preferred Qualifications 

  • Industry background: experience within a systems integrator (SI) or managed service provider (MSP) environment. 
  • Cluster management: hands-on experience with NVIDIA Base Command Manager and/or NVIDIA Mission Control. 
  • Network integration: understanding of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2) and how they interact with host PCIe and NVLink topologies. 
  • Container platforms: Kubernetes and Red Hat OpenShift installation and administration. 
  • Cloud architecture: AI-oriented compute design on AWS, Azure, Google Cloud, or OCI. 
  • AI and orchestration tooling: exposure to Rafay, NVIDIA Run:ai, NVIDIA Omniverse, Red Hat OpenShift AI, and Kubeflow. 

Certain states and localities require employers to post a reasonable estimate of salary range. A reasonable estimate of the current base pay range for this position is $110,000.00 to $150,000.00 annually. Actual salary will be based on a variety of factors, including shift, location, experience, skill set, performance, licensure and certification, and business needs. The range for this position in other geographic locations may differ. Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that is not included in the base pay.


The well-being of WWT employees is essential. When it comes to our benefits package, WWT has one of the best. We offer the following benefits to all full-time employees:

  • Health and Wellbeing: Health (Medical & Prescription), Dental, and Vision Care, Onsite Health Centers (MO & IL), Employee Assistance Program, Wellness program
  • Financial Benefits: Competitive Pay, Profit Sharing, 401k Plan with Company Matching, Life and Disability Insurance, Flexible Spending Accounts, Tuition Reimbursement
  • Paid Time Off: PTO & Holidays, Parental Leave, Medical Leave, Military Leave, Bereavement, Day of Caring
  • Additional Perks: Family Planning Benefits, Nursing Mothers Benefits, Voluntary Legal, Voluntary Supplemental Accident/Illness/Hospital, Voluntary ID Theft, Pet Insurance, Employee Discount Program

Note: This is not an all-encompassing list and should not be used as a complete description of the plan’s benefits. For more information, see our US benefits website at wwt.com/us-benefits.


We strive to create an environment where all employees are empowered to succeed based on their skills, performance, and dedication. Our goal is to cultivate a culture of belonging that encourages innovation, collaboration, and respect for all team members, ensuring that WWT remains a great place to work for all!


If you require accessibility accommodation(s) or adjustment during any stage of the hiring process, please let your WWT Recruiter know. The recruiter will work with you to understand your needs and help ensure an accessible experience throughout the interview process.


World Wide Technology is an Equal Opportunity Employer.


If you have any questions or concerns about this posting, please email [email protected]


#LI-AF1

#LI-Remote

Preferred Qualifications

World Wide Technology (WWT) strives to make a new world happen. WWT's work benefits clients and partners as much as it does its people and community across the globe.


Founded in 1990, WWT brings together strategy, deep technical expertise and world-class partnerships to help public and private sector organizations design, build and scale intelligent AI, digital, cybersecurity, cloud and infrastructure solutions. Through its Advanced Technology Center (ATC)—a collaborative ecosystem featuring state-of-the-art hardware and software—WWT enables clients and partners to conceptualize, test and validate innovative technology and then deploy solutions at scale using its global integration and distributions capabilities.


With more than 14,000 team members and over 60 locations globally, WWT's culture—grounded in core values and leadership philosophies—has been recognized by Fortune® and Great Place to Work for its commitment to innovation, trust and creating a great place to work for all. WWT provides products and services to large enterprise, global service provider and public sector clients in up to 130 countries across six continents. Softchoice, a World Wide Technology company, supports commercial and SMB markets in the U.S. and Canada.


Want to work with highly motivated individuals on high-performance teams? Join WWT today!


What is the Solutions Consulting & Engineering Team and why join?


Solutions Consulting & Engineering is an organization that is customer-focused and solutions-led. We deliver end-to-end and emerging solutions to drive customer satisfaction and increase profitability and growth. Our world-class management consulting, delivery excellence, and engineering brilliance enable our success. We embody the OneWWT mindset by bringing the right talent at the right time from anywhere within WWT to solve our customer’s problems. Our goal is to bring together business acumen with full-stack technical know-how to develop innovative solutions for our clients’ most complex challenges. 


About the Role 

As Domain Architect - AI Compute, you will be the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across a diverse portfolio of client environments, bridging the gap between architectural design and hands-on execution. You are a builder as much as an advisor: as comfortable configuring a cluster from the CLI as you are explaining that configuration to a C-level audience. 

 

As a global systems integrator, we don't simply operate static cloud environments. We design and deliver purpose-built, high-scale AI factories for some of the world's leading enterprises. In this role you will define the reference standard for compute infrastructure, moving beyond single-server administration to architect repeatable, scalable, and automated compute fabrics. You will act as technical lead on NVIDIA Cloud Partner (NCP) and private enterprise AI cloud deployments, owning the compute layer of the compute, network, and storage stack. 

 

Your time will be split roughly 60/40 between delivering complex AI infrastructure (60%) and providing pre-sales subject matter expertise (40%). You will lead the physical provisioning of NVIDIA DGX SuperPOD, NVIDIA DGX BasePOD, and Cisco AI POD environments, ensuring clients inherit platforms that are genuinely ready for day-2 operations, while helping the sales team scope and cost future deployments. 

 

Key Responsibilities 

Delivery and implementation  

Bare-metal build and provisioning 

  • Lead the physical provisioning of GPU platforms and clusters, including NVIDIA GB200/GB300 NVL72 rack-scale systems, NVIDIA DGX SuperPOD and DGX BasePOD reference architectures, HGX- and MGX-based systems, and Cisco AI PODs. 
  • Use NVIDIA Base Command Manager (BCM) for cluster provisioning, diskless boot, image management, and firmware lifecycle management. 
  • Harden and baseline the host operating system across the fleet. 
  • Establish cluster monitoring and lifecycle management with NVIDIA Mission Control. 
  • Build and execute zero-touch provisioning (ZTP) workflows that turn bare-metal hardware into production-ready nodes. 

 

Scheduler and workload configuration 

  • Define and enforce fair-share policies, fractional GPU allocation using Multi-Instance GPU (MIG), and preemption logic for multi-tenant environments. 
  • Ensure orchestration layers respect hardware topology, including NUMA affinity and PCIe topology, to protect performance. 
  • Implement and tune advanced schedulers: Slurm on bare metal, and NVIDIA Run:ai, Kueue, or Volcano on Kubernetes. 

 

Orchestration and day-2 operations 

  • Deploy and configure multi-cluster management and observability platforms such as Rafay. 
  • Implement high-fidelity telemetry using NVIDIA Data Center GPU Manager (DCGM) to monitor GPU health, thermal throttling, and Xid error rates. 
  • Lead the transition to day-2 operations, ensuring the environment is fully integrated with the client's identity providers and storage backends before handoff. 

 

Performance engineering 

  • Conduct acceptance and validation testing using NCCL tests and HPL to verify cluster performance against expected baselines. 
  • Perform kernel and OS-level tuning (hugepages, sysctl, and driver parameters) to optimize for high-bandwidth InfiniBand and RoCEv2 fabrics. 

 

Pre-sales SME and consulting 

Technical scoping and estimation 

  • Support the sales team by validating customer technical requirements and producing accurate level-of-effort (LOE) estimates for statements of work. 
  • Define the standard operating environment (SOE) used in proposals to ensure repeatability across engagements. 

 

Bill of materials validation and architecture 

  • Own the technical accuracy of the compute bill of materials. 
  • Verify that all components, including memory, NVMe storage, NICs, and optical transceivers, are aligned with NVIDIA-Certified Systems requirements and the qualified component lists in the relevant DGX, HGX, MGX, and NVL72 reference architectures. 
  • Match GPU platform selection to the client's actual workload profile and growth expectations. 

 

Client workshops 

  • Lead technical discovery workshops to determine workload requirements, distinguishing between the very different infrastructure profiles of traditional machine learning, LLM training, LLM inference, and Omniverse digital twin and visualization workloads. 

Skills Required

  • 10+ years of experience in HPC, AI, or data center infrastructure engineering
  • 7+ years of experience in a customer-facing architecture role
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • Deep working knowledge of NVIDIA GPU platforms across Hopper and Blackwell architectures, including Grace Hopper and Grace Blackwell systems
  • Knowledge of CUDA, cuDNN, and NCCL
  • Strong Linux systems engineering skills, including Ubuntu, RHEL, kernel tuning, driver management, and system hardening
  • Automation and infrastructure-as-code experience with Ansible and Python
  • Familiarity with Terraform or similar declarative tooling
  • Ability to translate technical architecture into commercial outcomes for senior stakeholders
  • Experience in a systems integrator or managed service provider environment
  • Experience with NVIDIA Base Command Manager and/or NVIDIA Mission Control
  • Understanding of InfiniBand NDR/HDR, RoCEv2, PCIe, and NVLink topologies
  • Experience installing and administering Kubernetes and Red Hat OpenShift
  • AI compute architecture experience on AWS, Azure, Google Cloud, or OCI
  • Exposure to Rafay, NVIDIA Run:ai, NVIDIA Omniverse, Red Hat OpenShift AI, and Kubeflow

World Wide Technology Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about World Wide Technology and has not been reviewed or approved by World Wide Technology.

  • Healthcare Strength Health insurance is considered excellent and affordable, with bundled medical, dental, and vision coverage that reduces out-of-pocket costs. Wellness programs and an Employee Assistance Program further enhance access to care.
  • Leave & Time Off Breadth Generous paid time off grows with tenure, alongside paid sick days and holidays. Parental and military leave are also provided.
  • Retirement Support Financial benefits include profit sharing and a 401(k) with company matching, complemented by life and disability insurance and tuition reimbursement. These elements strengthen total rewards beyond base salary.

World Wide Technology Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Maryland Heights, MO
9,000 Employees
Year Founded: 1990

What We Do

World Wide Technology is a systems integrator, provides information technology and supply chain solutions. Fueled by creativity and ideation, World Wide Technology strives to accelerate our growth and nurture future innovation. From our world class culture, to our generous benefits, to developing cutting edge technology solutions, WWT constantly works towards its mission of creating a profitable growth company that is a great place to work. We encourage our employees to embrace collaboration, get creative and think outside the box when it comes to delivering some of the most advanced technology solutions for our customers. At a glance, WWT was founded in 1990 in St. Louis, Missouri. We employ over 9,000 individuals and closed nearly $17 Billion in revenue. We have an inclusive culture and believe our core values are the key to company and employee success. WWT is proud to announce that it has been named on the FORTUNE "100 Best Places to Work For®" list for the 12th consecutive year!

Why Work With Us

Our extensive partnership with best in class technology companies, coupled with our strong culture allow for world class delivery of transformative business solutions driven by IT.

Similar Jobs

Softchoice Logo Softchoice

Architect

Cloud • Information Technology • Productivity • Security • Software
Remote
United States
110K-150K Annually

Datadog Logo Datadog

Senior Sales Engineer

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
8 Locations
6500 Employees
149K-198K Annually

T-Mobile Logo T-Mobile

Account Executive

Other • Utilities
In-Office or Remote
2 Locations
89016 Employees
43K-78K Annually

T-Mobile Logo T-Mobile

Account Executive

Other • Utilities
Remote
Iowa, USA
89016 Employees
43K-78K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account