Manager, Performance Research and Analysis

Posted One Month Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Lead performance strategy, characterization, and optimization for large-scale NVIDIA AI GPU clusters. Focus on RDMA, collective communications (NCCL/MPI), congestion control, load balancing, DPUs, storage for inference, and cluster observability via telemetry and Grafana dashboards. Perform deep RCA and coordinate mitigation across hardware, firmware, and software teams.
Summary Generated by Built In

NVIDIA is seeking a highly skilled and versatile Performance Research and Analysis Manager to join our Performance Group. This role will drive end-to-end performance strategy and execution for next-generation NVIDIA data centers and solutions based on GPU systems, NIC, Switch, DPU and Networking technologies. The ideal candidate will oversee, evaluating, and optimizing end-to-end AI GPU cluster-level performance for scaling out large scale distributed training and inference jobs communication. The role will focus heavily on RDMA, Networking Protocols, Collective Communication, Congestion Control, and Load Balancing algorithms. Secondarily, you will lead NVIDIA DPUs and Storage technologies for N-S use cases to support AI Inference jobs. Third, you will drive our Performance Dashboards and Observability for cluster-level performance analysis from a stream line telemetry across NICs, Switches, GPUs, and NVlink.

What you'll be doing:

  • Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation NVIDIA AI GPU clusters, focusing on large-scale distributed training and inference workloads.

  • Deeply evaluate and optimize NVIDIA Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.

  • Work on performance research and analysis of NVIDIA DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.

  • Drive the strategy for performance observability and dashboards across next-generation NVIDIA data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.

  • Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.

What we need to see:

  • B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.

  • 8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.

  • 3+ years of experience as an engineering team manager leading technical performance or R&D teams.

  • Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).

  • Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.

  • Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.

Ways to stand out from the crowd:

  • Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.

  • Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.

  • Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.

  • Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

#LI-Hybrid

Skills Required

  • B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience
  • 8+ years overall experience with deep expertise in High Performance Networking, RDMA, and systems-level performance
  • 3+ years experience as an engineering team manager leading technical performance or R&D teams
  • Hands-on experience analyzing and optimizing collective communication (NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads
  • Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization
  • Exceptional cross-team leadership, analytical thinking, and communication skills
  • Proven track record optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms for multi-thousand GPU deployments
  • Deep experience tuning advanced network mechanisms (adaptive routing, PFC/ECN, packet-spraying)
  • Experience building autonomous performance tools, AI-assisted RCA agents, or automated regression frameworks
  • Hands-on experience developing custom Grafana plugins or complex panels and using PromQL/LogQL in hyperscale/HPC environments

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

HiBob Logo HiBob

Senior AI Ops Engineer

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
IL
1350 Employees

HiBob Logo HiBob

Team Lead

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
IL
1350 Employees

HiBob Logo HiBob

Junior Field Deployment Engineer (FDE)

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
IL
1350 Employees

Akamai Technologies Logo Akamai Technologies

Senior C++ Low Level Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account