Research Intern

Posted 2 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Internship
Artificial Intelligence • Software
We use AI to understand human ability and match talent with the opportunities they're best suited for.
The Role
Conduct research on post-training, reinforcement learning with verifiable rewards, data generation, and language model evaluation. Design controlled experiments, develop benchmarks and scoring frameworks, analyze model failures, build research tooling and data pipelines, and communicate findings through reports and presentations. Collaborate with research scientists, engineers, AI teams, and domain experts, contributing to publications, benchmark releases, and frontier AI systems.
Summary Generated by Built In
About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Role

As a Research Scientist Intern at Mercor, you’ll work on research at the frontier of post-training, reinforcement learning with verifiable rewards (RLVR), data generation, and model evaluation.

You’ll investigate how datasets, rewards, and training methods affect the capabilities and behavior of large language models. This may include designing controlled experiments, developing new evaluation methodologies, conducting systematic failure analysis, and testing approaches to improve tool use, agentic behavior, and real-world reasoning.

You’ll work closely with research scientists, research engineers, and domain experts to turn open-ended questions into rigorous experiments. Your work will contribute to Mercor’s research agenda and may support external publications, benchmark releases, and the development of frontier AI systems.

What You’ll Do
  • Develop and investigate research questions related to post-training, RLVR, data quality, and model evaluation.

  • Design and run controlled experiments to understand how datasets, rewards, and training strategies affect model performance.

  • Study reward-shaping and post-training methods, including approaches such as GRPO and DAPO.

  • Develop methods for measuring data quality, usability, and performance uplift on key benchmarks.

  • Design and evaluate datasets, rubrics, evaluators, and scoring frameworks for complex model capabilities.

  • Conduct systematic error analysis to identify model failure modes and opportunities for improvement.

  • Analyze experimental results and communicate findings through clear reports, research artifacts, and presentations.

  • Build the research tooling and data pipelines needed to conduct experiments at scale.

  • Collaborate with research scientists, research engineers, applied AI teams, and domain experts producing training and evaluation data.

  • Contribute to research publications, benchmark releases, and other public research outputs where appropriate.

What We’re Looking For
  • Currently pursuing a master’s or PhD in computer science, machine learning, statistics, mathematics, or another relevant field.

  • Demonstrated ability to formulate research questions, design experiments, and draw sound conclusions from empirical results.

  • Demonstrated experience in at least one of the following:

    • Training, fine-tuning, or evaluating language models.

    • Agentic AI system, RL environments

    • Developing benchmarks, evaluation methodologies, or data-quality measures.

  • At least one publication or open source project.

  • Strong programming skills, particularly in Python, and the ability to write reliable research code.

  • Familiarity with machine learning fundamentals, experimental design, and statistical analysis.

  • Intellectual curiosity.

  • Comfort operating in a fast-paced research environment with rapid iteration and a high degree of ownership.

Nice to Have
  • Previous research experience in language models, reinforcement learning, model evaluation, or post-training.

  • Experience training, fine-tuning, or evaluating language models.

  • Familiarity with RLVR techniques, reward modeling, or agentic AI systems.

  • Experience developing benchmarks, evaluation methodologies, or data-quality measures.

  • Research publications or submissions at competitive CS conferences such as ACL, NeurIPS, ICML, ICLR, or EMNLP.

  • Research papers, technical reports, open-source projects, or other work samples demonstrating relevant skills.

Why Mercor

Impact: Your work powers how AI labs train and deploy their models

Learning: Get early exposure to frontier AI research and engineering

Growth: Work with a high-velocity team where interns ship to production

Benefits

  • Mentorship from experienced researchers.

  • Work on real, high-impact projects.

  • $1.5K monthly stipend for meals

  • $200 monthly laundry reimbursement

  • $200 monthly personal wellness reimbursement

  • Free Equinox membership

  • Team events and offsites.

  • Potential full-time return offer.

Skills Required

  • Currently pursuing a master's or PhD in computer science, machine learning, statistics, mathematics, or a related field
  • Ability to formulate research questions, design experiments, and draw sound conclusions from empirical results
  • Experience training, fine-tuning, or evaluating language models; developing agentic AI systems or reinforcement learning environments; or developing benchmarks, evaluation methodologies, or data-quality measures
  • At least one publication or open source project
  • Strong programming skills, particularly in Python, and ability to write reliable research code
  • Familiarity with machine learning fundamentals, experimental design, and statistical analysis
  • Intellectual curiosity
  • Comfort operating in a fast-paced research environment with rapid iteration and high ownership

Mercor Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Mercor and has not been reviewed or approved by Mercor.

  • Fair & Transparent Compensation — Pay is considered competitive across many roles, with clear hourly ranges and an hourly/pay‑per‑task mix designed to align rates with expertise. The structure emphasizes transparent, appropriate pay levels and guarantees payment for legitimate logged time.
  • Strong & Reliable Incentives — Payments are processed on a predictable weekly cadence via Stripe/Wise, and some tracks offer additional weekly bonus incentives for top performers. This combination of regular payouts and performance bonuses supports dependable earnings when projects are active.
  • Equity Value & Accessibility — Select full‑time roles include generous equity grants alongside cash perks such as relocation and housing bonuses. These elements increase total compensation for those positions.

Mercor Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
2,217 Employees
Year Founded: 2023

What We Do

We use AI to understand human ability and match talent with the opportunities they're best suited for.

Similar Jobs

Coinbase Logo Coinbase

User Research Intern

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Hybrid
San Francisco, CA, USA
4700 Employees
60-60 Annually

SK hynix Logo SK hynix

Research Intern - Loop Engineering

Information Technology • Semiconductor • Industrial
In-Office
San Jose, CA, USA
328 Employees
35-45 Hourly
In-Office
San Diego, CA, USA
1001 Employees
18-24 Hourly

Bobyard Logo Bobyard

Computer Vision Research Engineer - Intern

Information Technology • Internet of Things
In-Office
San Francisco, CA, USA
30 Employees
55-55 Hourly

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account