Mercor Research Fellowship — APEX

Posted 4 Hours Ago
Be an Early Applicant
3 Locations
In-Office or Remote
40K-80K Annually
Entry level
Artificial Intelligence • Software
We use AI to understand human ability and match talent with the opportunities they're best suited for.
The Role
Design, build, and validate new AI benchmarks and evaluation methodologies. Fellows create task specifications and grading rubrics, collaborate with domain experts, run frontier models, analyze failures, test for contamination and gaming, and publish findings through papers, datasets, leaderboards, or internal methodologies. The fellowship provides mentorship, compute, expert labor, and access to enterprise evaluation problems.
Summary Generated by Built In
About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Fellowship

Mercor’s APEX benchmark family measures whether frontier AI models can actually do economically valuable work: multi-hour agentic tasks in investment banking and corporate law, real professional accounting workflows, real-world software engineering, and graduate-level science. Every APEX benchmark is built and validated with Mercor’s network of domain experts — not written from a textbook.

The Mercor Research Fellowship funds people to build the next generation of benchmarks and evaluation techniques. You pitch a benchmark or eval methodology you want to build — a new domain, a harder task format, a better way to measure agentic reliability — and if selected, you get the time, compute, expert labor, and mentorship to design, implement, and release it end to end.

You’ll work directly with the APEX research team, get access to real enterprise evaluation problems from Mercor’s Fortune 500 and frontier-lab partners, and see your benchmark shape how the industry measures AI capability.

Program Details

  • Duration: 3–6 months, rolling admission

  • Commitment: minimum 30 hours/week; full-time preferred

  • Location: remote, or in-person at Mercor’s San Francisco office

  • Admission: apply with a specific benchmark or eval technique you want to build — the fellowship is funded around your pitch, not a generic research rotation

What You’ll Do

  • Propose and scope a new benchmark or evaluation technique in a domain APEX doesn’t yet cover well, or a meaningfully harder version of one it does.

  • Design task specifications and grading rubrics in partnership with Mercor’s network of vetted domain experts — lawyers, accountants, engineers, scientists, and consultants.

  • Build and validate the benchmark: pilot tasks, calibrate scoring, and stress-test for contamination and gameable shortcuts.

  • Run frontier models against your benchmark and analyze where and why they fail.

  • Publish your results — as a paper, an open dataset, a new leaderboard on APEX, or a methodology the APEX team adopts internally.

  • Partner with Mercor’s research and engineering teams to fold what you learn back into APEX’s public benchmark family.

Focus Areas

  • Long-horizon, multi-app agentic tasks in professional services (law, finance, consulting) — extending APEX-Agents

  • Real-world software engineering evaluation beyond issue resolution — extending APEX-SWE

  • Professional accounting and finance workflows — extending APEX-Accounting

  • AI-for-Science evals: research-level mathematics, biology, materials science, and theoretical physics

  • Novel evaluation methodology: contamination resistance, rubric design, human-vs-model grading agreement, cost-adjusted scoring

  • Strong pitches outside this list are welcome — we fund the best ideas, not the closest fit to a template.

What We’re Looking For

  • Genuine interest in evaluation as a research discipline — not just a stepping stone to a model-building role.

  • Background in CS, ML, statistics, or an adjacent field (measurement, psychometrics, HCI, social science); no requirement to have published in ML venues.

  • A specific, well-scoped idea for a benchmark or eval technique you want to build — the fellowship is built around your pitch.

  • Comfortable in a startup environment: fast iteration, direct access to real customer problems, less hand-holding than an academic lab.

  • Able to commit at least 20 hours/week for the duration of the fellowship — during a leave, over a summer, or a flexible stretch of a PhD.

  • Bonus: experience with agentic evaluation, RL environments, or domain expertise in law, finance, medicine, or a scientific field.

Compensation & Benefits

  • 3 month stipend of $40,000 or 6 month stipend of $80,000

  • Unlimited API credits, plus a dedicated budget for GPU compute and paid expert/human-data time

  • Weekly 1:1 mentorship with a member of the APEX research team, plus regular access to the broader research org

  • Access to frontier model APIs, Mercor’s internal evaluation infrastructure, and — where appropriate — real enterprise evaluation problems from Mercor’s customers

  • Optional desk in Mercor’s San Francisco office for fellows who want to be in person

  • Introductions to Mercor’s network of researchers across frontier labs and academia

  • Standout fellows are considered for a full-time offer on the APEX research team at the end of the fellowship

Skills Required

  • Background in computer science, machine learning, statistics, measurement, psychometrics, human-computer interaction, social science, or an adjacent field
  • Specific, well-scoped idea for a benchmark or evaluation technique to build
  • Ability to commit at least 20 hours per week for the fellowship duration
  • Comfort working in a fast-paced startup environment with limited hand-holding
  • Genuine interest in evaluation as a research discipline
  • Experience with agentic evaluation, reinforcement learning environments, or domain expertise in law, finance, medicine, or a scientific field

Mercor Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Mercor and has not been reviewed or approved by Mercor.

  • Fair & Transparent Compensation Pay is considered competitive across many roles, with clear hourly ranges and an hourly/pay‑per‑task mix designed to align rates with expertise. The structure emphasizes transparent, appropriate pay levels and guarantees payment for legitimate logged time.
  • Strong & Reliable Incentives Payments are processed on a predictable weekly cadence via Stripe/Wise, and some tracks offer additional weekly bonus incentives for top performers. This combination of regular payouts and performance bonuses supports dependable earnings when projects are active.
  • Equity Value & Accessibility Select full‑time roles include generous equity grants alongside cash perks such as relocation and housing bonuses. These elements increase total compensation for those positions.

Mercor Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
2,217 Employees
Year Founded: 2023

What We Do

We use AI to understand human ability and match talent with the opportunities they're best suited for.

Similar Jobs

Byrd's House Detective DBA BHD Home and Mold Inspections Logo Byrd's House Detective DBA BHD Home and Mold Inspections

Virtual Assistant

Agency • Logistics • Real Estate • Virtual Reality • Consulting • Hospitality • Data Privacy
Remote or Hybrid
United States
100 Employees
35-45 Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Senior Casualty Claims Specialist, Attorney Represented - Northeast

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
7 Locations
40000 Employees
61K-126K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
4 Locations
55000 Employees
245K-335K Annually

Capital One Logo Capital One

Sr. Director, Technical Program Management (Remote-Eligible)

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
5 Locations
55000 Employees
245K-336K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account