Senior Solutions Architect, HPC and AI

Posted 25 Days Ago
Be an Early Applicant
3 Locations
In-Office
184K-357K
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
As a Senior Solutions Architect, HPC and AI, you will validate and debug large-scale GPU clusters, assist customers with technical issues, and contribute to deep learning advancements.
Summary Generated by Built In

NVIDIA is looking for a Field Escalation Solution Architect with experience in validation and debugging of large-scale GPU clusters focused on performance. As part of the Solution Architecture organization, we work with the most sophisticated computing hardware and software, driving the latest deep learning and machine learning breakthroughs with NVIDIA’s enterprise customers. This role offers an excellent opportunity to build your career in the rapidly growing field of deep learning while enabling the world's most successful technology companies. Primary responsibilities will be to validate and debug customer cluster performance issues, functional bottlenecks and drive customer technical engagements around NVIDIA products and technologies. Join us in this exciting endeavor!

What you’ll be doing:

  • A considerable part of the day-to-day job is staying up to date on pioneering High Performance Computing, Deep Learning and Machine Learning ecosystems. You'll be called on to help architect and scale high-performance, distributed AI infrastructure on-prem or in the cloud built with the latest NVIDIA GPU supercomputers for new and existing customers.

  • Address and resolve problems starting from the bare metal level, all the way up to the operating system, software stack, and application level.

  • Share knowledge with different teams by delivering demos, assisting with proof-of-concepts, and writing papers and developer blogs. By collaborating with executives and engineering, address sophisticated problems and help bring NVIDIA's premiere technologies to life in the cloud and in the datacenter.

  • Work directly with developers and hardware architects to debug cluster performance issues, identify new requirements, and improve workflows.

  • Will be engaged by the account team when extra analysis is required in debugging customer issues.

  • Provide additional expertise to enable the account team to be more adaptable to the customer and product engineering to get more actionable data at speed of light making them more efficient.

  • Building custom product demonstrations and POCs for solutions that address critical business needs of our customers.

What we need to see:

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or other Engineering fields or equivalent experience.

  • 8+ years of work-related experience in NVIDIA and/or accelerated computing technologies.

  • Platform level understanding of server architecture, PCIe topology, GPUs, NICs, Linux OS and kernel drivers.

  • Networking experience, including knowledge of Ethernet, InfiniBand or other networking protocols.

  • Experience working with DevOps on-prem or in cloud environments, including but not limited to Docker/Containers, cloud APIs, IaaS and Data Center deployments.

  • SLURM, Kubernetes, and/or other job scheduler use, deployment, and debugging skills.

  • Deep understanding of dense data center design, including computing, storage, networking, cloud APIs, and IaaS.

  • Effective time management and capable of balancing multiple tasks.

  • Strong analytical and problem-solving skills.

  • Strong communication skills, both written and verbal, with the ability to collaborate and coordinate efficiently across multi-functional teams in engineering, sales, marketing, product, and program management.

Ways to stand out from the crowd:

  • Demonstrated Communication Collectives (NCCL) experience.

  • Excellent customer-facing skills and background.

  • Platform design engineering, coding and proficient debugging skills including experience in C/C++, Linux kernel, virtualization and drivers, profilers/performance analysis tools (NSys).

  • Familiarity with NVIDIA systems/SDKs (e.g. CUDA), NVIDIA Networking technologies (e.g., RoCE, InfiniBand), Switch interconnects and/or ARM CPU solutions through hands-on experience.

  • Understanding of Deep Learning and Machine Learning frameworks (TensorFlow or PyTorch), LLM, MLOps, DevOps, and workflows applying cloud technologies, using Docker/containers, Kubernetes, cloud APIs, and data center deployments, among others.

We make extensive use of conferencing tools, but occasional travel (25%) is required for a local on-site visit to customers and data science conferences.

With highly competitive salaries, a comprehensive benefits package, and an excellent engineering culture, NVIDIA is widely considered to be one of the technology industry's most desirable employers. NVIDIA has some of the most innovative people working on significant problems that define the field of ML/DL, data science, and graphics.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 2, 2025.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

Cuda
Docker
Ethernet
Infiniband
Kubernetes
Linux
Nvidia Gpu
Slurm
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Spectrum Logo Spectrum

Manager, Web Development - Spectrum Reach

Information Technology • Internet of Things • Mobile • On-Demand • Software
In-Office
Dallas, TX, USA

Spectrum Logo Spectrum

Account Executive

Information Technology • Internet of Things • Mobile • On-Demand • Software
In-Office
Austin, TX, USA

Pfizer Logo Pfizer

I&I Gastroenterology Area Business Manager - Houston, TX

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
2 Locations
133K-268K Annually

Capital One Logo Capital One

Work at Home COAF Ops Sr. Coordinator

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
2 Locations
50K-50K Annually

Similar Companies Hiring

Scrunch AI Thumbnail
Software • SEO • Marketing Tech • Information Technology • Artificial Intelligence
Salt Lake City, Utah
Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Standard Template Labs Thumbnail
Software • Information Technology • Artificial Intelligence
New York, NY
10 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account