Product Manager, APEX

Posted 7 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
180K-300K Annually
Senior level
Artificial Intelligence • Software
We use AI to understand human ability and match talent with the opportunities they're best suited for.
The Role
Own Mercor’s APEX AI evaluation product, including roadmap, benchmark portfolio, public and private leaderboards, evaluation integrity, and end-to-end infrastructure. Partner with research, engineering, operations, GTM teams, frontier AI labs, and enterprise customers to develop trusted benchmarks, communicate results, drive adoption, and connect evaluation demand to business impact. Write product specifications, analyze evaluation failures, and improve platform capabilities hands-on.
Summary Generated by Built In
About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Role

Mercor’s AI Productivity Index (APEX) assesses how effectively frontier AI models can perform economically valuable work. We're looking for a Product Manager to own and scale the APEX brand and public leaderboard. In this role, you will define the positioning, strategy, roadmap, and operational excellence of Mercor's evaluation products, benchmarks, and public and private leaderboards. You will work across Research, Engineering, Operations, and Go-To-Market teams to transform evaluation datasets into trusted industry benchmarks that influence model development and purchasing decisions across the AI ecosystem.

You will serve as the product owner for APEX and our eval platform, driving benchmark innovation, evaluation integrity, infrastructure, customer adoption, and business impact. You will work directly with frontier AI labs and enterprise customers, representing Mercor as a thought leader in AI evaluation and measurement. You will partner with Mercor’s world-class benchmark research team, which includes the first authors from many popular benchmarks including Tau Bench, SciCode, PostTrainBench, and more.

The ideal candidate combines strong product judgment, technical fluency, operational rigor, and customer-facing experience, with a passion for turning emerging model capabilities into a credible measure of what AI can actually do for the economy.

What You'll Do
  • Own the roadmap and portfolio strategy: Set and influence priorities across new benchmark development, leaderboard launches, infrastructure investment, and expansion into new evaluation categories. Decide which domains earn a slot, when a benchmark has saturated, and what replaces it. Identify opportunity areas for strategic partnership.

  • Run the intake for new benchmarks: Evaluate and prioritize proposals from research, customers, and GTM against real demand and company strategy.

  • Enforce eval integrity: Own contamination policy, holdout strategy, versioning, auditability, and release cadence. Publish methodology clearly enough that a skeptical researcher can reconstruct our results.

  • Build end-to-end eval pipeline: Own the path from eval run to published result, including harness execution, grading, model onboarding, hyperparameter scaffolds, cost and latency reporting, and the leaderboard surface itself.

  • Work directly with labs and customers: Understand evaluation needs, present and defend results, and turn what you hear into the next benchmark. Support GTM on launches, partnerships, and thought leadership to influence customer model development strategies.

  • Close the loop to the business: Connect leaderboard demand signals to loss analysis investments and dataset production. Track adoption, usage, and downstream revenue to continuously justify leaderboard ROI.

  • Get in the weeds: Write specs and PRDs, but also read trajectories, spot-check failures, and make small PRs to unblock yourself and continuously improve the system.

What We're Looking For
  • Experience: 5+ years in product management, technical program management, or a customer-facing technical role. Prior background in SWE, ML, or DS strongly preferred.

  • Eval literacy: You reason fluently about rubric design, inter-rater reliability, agentic harnesses, contamination and overfitting, and what a small sample can and can't support. You can tell a real capability gap from measurement noise.

  • Public judgment: You'll publish numbers about other people's models. You know how to be neutral, precise, and defensible under scrutiny.

  • Taste for what matters: Strong instincts for which capabilities are actually worth measuring, and what results are actually worth highlighting.

  • Entrepreneurial: A track record of building a product, program, or business line from nothing.

  • High Ownership: You take full accountability for outcomes, not just outputs.

  • Independence: Able to self-direct in ambiguous contexts, creating clarity for others.

  • Stakeholder alignment: You can bring research, engineering, operations, and GTM around a shared methodology.

Benefits

  • Bi-annual performance bonus structure

  • Generous equity grant vested over 4 years

  • Up to $15k Relocation bonus

  • $10K housing bonus (if you live within 0.5 miles of our office)

  • $1.5K monthly stipend for meals

  • Free Equinox membership

  • $200 monthly laundry reimbursement

  • $200 monthly personal wellness reimbursement

  • Health, Dental, Vision insurance

Skills Required

  • 5+ years of experience in product management, technical program management, or a customer-facing technical role
  • Fluency in rubric design, inter-rater reliability, agentic harnesses, contamination, overfitting, and evaluation methodology
  • Strong judgment regarding AI capabilities, benchmark relevance, and public reporting of model results
  • Track record of building a product, program, or business line from nothing
  • Ability to take full accountability for outcomes
  • Ability to self-direct in ambiguous contexts and create clarity for others
  • Ability to align research, engineering, operations, and go-to-market stakeholders around shared methodology
  • Prior software engineering, machine learning, or data science background
  • Technical fluency, operational rigor, and customer-facing experience

Mercor Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Mercor and has not been reviewed or approved by Mercor.

  • Fair & Transparent Compensation Pay is considered competitive across many roles, with clear hourly ranges and an hourly/pay‑per‑task mix designed to align rates with expertise. The structure emphasizes transparent, appropriate pay levels and guarantees payment for legitimate logged time.
  • Strong & Reliable Incentives Payments are processed on a predictable weekly cadence via Stripe/Wise, and some tracks offer additional weekly bonus incentives for top performers. This combination of regular payouts and performance bonuses supports dependable earnings when projects are active.
  • Equity Value & Accessibility Select full‑time roles include generous equity grants alongside cash perks such as relocation and housing bonuses. These elements increase total compensation for those positions.

Mercor Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
2,217 Employees
Year Founded: 2023

What We Do

We use AI to understand human ability and match talent with the opportunities they're best suited for.

Similar Jobs

Boeing Logo Boeing

Property Management Specialist - Millennium Space Systems

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
El Segundo, CA, USA
170000 Employees
63K-93K Annually

Boeing Logo Boeing

Sales Representative

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
El Segundo, CA, USA
170000 Employees
129K-175K Annually

Dynatrace Logo Dynatrace

Operations Specialist

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
United States
5600 Employees
116K-145K Annually

Boeing Logo Boeing

C-17 Associate Mechanical System Design & Analysis Engineer

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Long Beach, CA, USA
170000 Employees
91K-123K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account