Senior HPC and AI Networking Performance Research and Analysis Engineer

Posted 17 Days Ago
Be an Early Applicant
Santa Clara, CA
Senior level
Artificial Intelligence • Hardware • Robotics • Software • Metaverse
The Role
This role involves profiling and analyzing AI workloads on large-scale GPU and CPU clusters for distributed deep learning, focusing on networking performance. Responsibilities include benchmarking, identifying performance bottlenecks, developing analysis tools, and collaborating across hardware and software teams to set performance expectations.
Summary Generated by Built In

Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines to learn, perceive, reason and solve problems. Today, visual computing is a crucial tool in helping people get along with technology, and NVIDIA has extended its technology into datacenters, mobile devices and cars. There has never been a more exciting time to join our team - if this role sounds like a fit for you, we'd love to hear from you!

NVIDIA is seeking a Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.

What you'll be doing:

  • Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on NVIDIA supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).

  • Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.

  • Implementing performance analysis tools.

  • Collaborating with many teams from hardware to software to provide performance analysis insights.

  • Defining performance test planning , setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

What we need to see:

  • B.Sc in Computer Science or Software Engineering or equivalent experience

  • 5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)

  • Demonstrated Performance Analysis skills and methodologies.

  • Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).

  • Fast and self-learning capabilities with strong analytical and problem-solving skills.

  • Programming Languages: Python, Bash and C languages

  • Experience with Linux OS distros.

  • Great teammate with good communication and interpersonal skills

Ways to stand out from the crowd:

  • In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.

  • Knowledge in CUDA, and NCCL libraries.

  • Knowledge in Congestion Control algorithms.

  • In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).

  • Strong Performance Analysis skills and methodologies using modern tools.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. We have a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world!

#LI-Hybrid

The base salary range is 148,000 USD - 276,000 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

Bash
C
Python
The Company
HQ: Santa Clara, CA
21,960 Employees
On-site Workplace
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Afterpay Logo Afterpay

Senior Data Platform Engineer

Fintech • Payments • Software • Financial Services
Hybrid
8 Locations
900 Employees
126K-223K Annually

Voltage Park Logo Voltage Park

Platform Engineer

Artificial Intelligence • Cloud • Hardware • Machine Learning • Other • Software • Infrastructure as a Service (IaaS)
San Francisco, CA, USA
51 Employees
120K-180K Annually

Crusoe Energy Systems Logo Crusoe Energy Systems

Electrical Engineer - Capital Projects

Cloud • Greentech • Other • Energy
Hybrid
San Francisco, CA, USA
450 Employees
170K-200K Annually

Crusoe Energy Systems Logo Crusoe Energy Systems

Electrical Engineer - Facilities

Cloud • Greentech • Other • Energy
Hybrid
San Francisco, CA, USA
450 Employees
170K-200K Annually

Similar Companies Hiring

Jobba Trade Technologies, Inc. Thumbnail
Software • Professional Services • Productivity • Information Technology • Cloud
Chicago, IL
45 Employees
RunPod Thumbnail
Software • Infrastructure as a Service (IaaS) • Cloud • Artificial Intelligence
Charlotte, North Carolina
53 Employees
Hedra Thumbnail
Software • News + Entertainment • Marketing Tech • Generative AI • Enterprise Web • Digital Media • Consumer Web
San Francisco, CA
14 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account