Product Evaluations Lead - Gen AI Software

Posted 14 Days Ago
Be an Early Applicant
Santa Clara, CA, USA
In-Office
224K-357K Annually
Expert/Leader
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Lead end-to-end evaluation and benchmarking of NVIDIA's Generative AI (Nemotron) models for strategic ISV partners. Design evaluation frameworks, experiments, and metrics; translate results into go/no-go recommendations; deliver dashboards and executive readouts; partner with research, product, and engineering to set release criteria and drive model-improvement signals and roadmap priorities.
Summary Generated by Built In

NVIDIA is seeking a highly analytical Product Evaluations Lead to own the evaluation, measurement, and go-forward analysis of our Generative AI Software including the Nemotron family as they are adopted by our most strategic enterprise ISV partners. As the evaluation's driver for our most strategic partners and products, you will help drive the experiments, benchmarks, and translate results into crisp insights and prioritized next steps that guide Product, Research, and leadership. This role operates at the intersection of rigorous data science, GenAI product strategy, and executive communication, staying ahead of the rapidly advancing reasoning and agentic AI landscape. You will directly influence model release criteria, product direction, and the roadmap for how NVIDIA's GenAI software is measured and improved!

What You Will Be Doing:

  • Drive end-to-end evaluation and analysis of NVIDIA's GenAI libraries, primarily Nemotron models for a focused set of strategic ISV partners, translating complex model behavior into clear go/no-go recommendations and prioritized next steps.

  • Develop and own rigorous evaluation frameworks, partner experimentation, and benchmarks that measure model quality, reasoning, and agentic capabilities against partner requirements and real-world enterprise use cases.

  • Share learnings and impact. Deliver regular readouts, dashboards, and progress tracking on evaluation results, adoption status, and recommended next steps to Product, Engineering, Research, and leadership stakeholders, consistent with NVIDIA's culture.

  • Provide analytical judgment under high uncertainty, balancing model quality, risk, and impact from incomplete or noisy signals to inform high-stakes release and integration decisions.

  • Partner with research scientists and product teams to translate emerging model capabilities into measurable release criteria, and identify the data and signals that feed the model-improvement flywheel.

  • Represent partner and product evaluation needs to internal teams. Contribute to the product roadmap by synthesizing cross-partner and cross-industry patterns captured from strategic engagements.

What we need to see:

  • 12+ years of experience in data science, model evaluation, experimentation, or analytics roles, with a focus on measuring AI/ML or GenAI systems.

  • Master's or PhD in a quantitative field (e.g., Data Science, Statistics, Computer Science, Economics) or equivalent experience.

  • Proven track record leading model evaluation and experimentation that directly informed high-stakes, go/no-go product or model-release decisions.

  • Strong hands-on skills in Python and statistical/experimental methods, including A/B testing, causal measurement (e.g., org-level holdouts), and metric design and failure analysis.

  • Direct experience evaluating and benchmarking LLMs or GenAI systems including reasoning, agentic, and/or multimodal capabilities, and building high-quality evaluation datasets and human-evaluation strategies.

  • Ability to translate complex, ambiguous model behavior into clear narratives, reporting, and recommendations for research, product, and executive audiences.

  • Proven ownership operating cross-functionally across research, engineering, product, and leadership to drive alignment and decisions in a high-velocity environment.

Ways to Stand Out from the Crowd:

  • Founding or early-team experience standing up a GenAI evaluation or analytics function defining north-star metrics and rapid-experimentation frameworks that accelerated and de-risked release decisions.

  • Track record of influencing complex model and product decisions through positive relationships, showing empathy for partner needs and an instinct for translating them into measurable improvements.

  • Hands-on experience with reward models, data flywheels, or evaluation signals that directly improved model performance.

  • Published or applied work in LLM capability evaluation or benchmarking (e.g., peer-reviewed research), paired with agility in a high-velocity GenAI landscape.

  • Highly collaborative standout colleague, able to build deep trust with engineers, researchers, executives, and multi-functional teams at both NVIDIA and partner organizations.

We're entering a defining chapter in NVIDIA's history! This role offers an outstanding opportunity to build how our Enterprise Generative AI Software is measured, trusted, and adopted by the world's leading partners. You'll have the support of Engineering, NVIDIA Research, Solutions Architects, and Product teams, and your analysis will directly influence which models ship and how our Enterprise AI business grows. This is an unparalleled opportunity for professional growth and impact in a high visibility team that’s at the forefront of innovation. 

NVIDIA is widely considered one of the technology world’s most desirable employers. We have some of the world's most forward-thinking and hardworking people on our team. If you're creative and autonomous, we want to hear from you! NVIDIA benefits is available online at Benefits and Support Programs | NVIDIA Benefits

#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 12, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Skills Required

  • 12+ years experience in data science, model evaluation, experimentation, or analytics with focus on measuring AI/ML or GenAI systems
  • Master's or PhD in a quantitative field (Data Science, Statistics, Computer Science, Economics) or equivalent experience
  • Proven track record leading model evaluation and experimentation that informed high-stakes go/no-go product or model-release decisions
  • Strong hands-on skills in Python and statistical/experimental methods, including A/B testing, causal measurement, metric design, and failure analysis
  • Direct experience evaluating and benchmarking LLMs or GenAI systems, including reasoning, agentic, and/or multimodal capabilities, and building evaluation datasets and human-evaluation strategies
  • Ability to translate complex, ambiguous model behavior into clear narratives, reporting, and recommendations for research, product, and executive audiences
  • Proven ownership operating cross-functionally across research, engineering, product, and leadership to drive alignment and decisions
  • Founding or early-team experience standing up a GenAI evaluation or analytics function and defining north-star metrics
  • Hands-on experience with reward models, data flywheels, or evaluation signals that directly improved model performance
  • Published or applied work in LLM capability evaluation or benchmarking (e.g., peer-reviewed research)
  • Highly collaborative, able to build trust with engineers, researchers, executives, and partner organizations

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Pfizer Logo Pfizer

Director, Site Management & Monitoring

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
2 Locations
121990 Employees
177K-294K Annually

Applied Systems Logo Applied Systems

Senior Product Manager

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Remote or Hybrid
United States
3079 Employees
100K-180K Annually

Applied Systems Logo Applied Systems

Data Engineer

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Remote or Hybrid
United States
3079 Employees
70K-120K Annually

SambaSafety Logo SambaSafety

Business Intelligence Analyst

Insurance • Logistics • Software • Transportation • Business Intelligence
Remote or Hybrid
United States
300 Employees
100K-110K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
Chicago, Illinois
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account