Research Engineer, Environments

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
In-Office
180K-500K Annually
Mid level
Artificial Intelligence • Software
We use AI to understand human ability and match talent with the opportunities they're best suited for.
The Role
Build end-to-end RL environments and verifiers from enterprise data: ship workflow extraction and grading models, automate task-refinement pipelines, platformize sandbox app clones with real data, deliver data to customers, and scale environment production while preserving high fidelity.
Summary Generated by Built In
About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Role

You’ll work with large enterprises to capture their data and transform it into high-fidelity RL environments for capability evaluations and training datasets for frontier labs. We focus on pushing the frontier of world-building, verifier engineering, and more alongside our partners.

Your goal will be to automate the process of building evals for real work in the economy.

What You'll Do
  • Ship models for workflow extraction, classification, and grading.

  • Engineer autonomous task refinement processes which distill data taste into pipelines.

  • Deliver data to customers and deploy into real engagements.

  • Help define the future of agentic transformation for enterprises around the world.

  • Deeply learn about the intricacies of enterprises through building evaluations for all aspects of work.

  • Build end-to-end environments for labs & enterprises by platformizing sandbox app clones, load real data into the sandboxes, build prompts from real workflows, and write verifiers leveraging enterprise expertise & golden outputs.

  • Systematize the production of environments to scale throughput while maintaining high-quality worlds and verifiers.

What We're Looking For
  • Prior experience shipping environments – you’ve contributed to an OSS framework, built environments at previous companies, or worked on agentic evaluations.

  • Strong full-stack engineering skills – you’ll be responsible for everything from infrastructure to app code to analytics

  • Bias to action – this team is focused on shipping evals, not just philosophizing about them.

  • Curiosity – being biased towards understanding and digging deep into model behavior and actually looking at the data.

  • Sweat the details that make a simulation indistinguishable from the real thing and have systems-level thinking skills that allow you to scale up quality.

Nice to Have
  • Experience with Temporal, Modal, or similar orchestration/compute services

  • Experience with synthetic data generation for frontier models.Past work auditing and scrutinizing industry-standard evaluations

Benefits
  • Semi-annual performance bonus structure

  • Generous equity grant vested over 4 years

  • Up to $15k Relocation bonus

  • $10K housing bonus (if you live within 0.5 miles of our office)

  • $1.5K monthly stipend for meals

  • Free Equinox membership

  • $200 monthly laundry reimbursement

  • $200 monthly personal wellness reimbursement

  • Health, Dental, Vision insurance

Skills Required

  • Prior experience shipping environments (OSS contributions, previously built environments, or agentic evaluations)
  • Strong full-stack engineering skills (infrastructure, app code, analytics)
  • Ability to platformize and scale environment production while maintaining quality
  • Bias to action and curiosity about model behavior and data
  • Systems-level thinking and attention to simulation detail
  • Work in-person five days a week in San Francisco, New York City, or London offices
  • Experience with Temporal, Modal, or similar orchestration/compute services
  • Experience with synthetic data generation for frontier models or auditing standard evaluations

Mercor Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Mercor and has not been reviewed or approved by Mercor.

  • Fair & Transparent Compensation Pay is considered competitive across many roles, with clear hourly ranges and an hourly/pay‑per‑task mix designed to align rates with expertise. The structure emphasizes transparent, appropriate pay levels and guarantees payment for legitimate logged time.
  • Strong & Reliable Incentives Payments are processed on a predictable weekly cadence via Stripe/Wise, and some tracks offer additional weekly bonus incentives for top performers. This combination of regular payouts and performance bonuses supports dependable earnings when projects are active.
  • Equity Value & Accessibility Select full‑time roles include generous equity grants alongside cash perks such as relocation and housing bonuses. These elements increase total compensation for those positions.

Mercor Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
2,217 Employees
Year Founded: 2023

What We Do

We use AI to understand human ability and match talent with the opportunities they're best suited for.

Similar Jobs

In-Office
San Francisco, CA, USA
2217 Employees

OpenAI Logo OpenAI

Research Engineer, Frontier Evals & Environments

Artificial Intelligence • Machine Learning • Generative AI
In-Office
San Francisco, CA, USA
4500 Employees
205K-380K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate II

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Carlsbad, CA, USA
16000 Employees
15-24 Hourly

Mastercard Logo Mastercard

Vice President, Product Management, Agentic Commerce

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
San Francisco, CA, USA
38800 Employees
204K-391K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account