RL Environments Engineer

Reposted 3 Days Ago
Be an Early Applicant
Mountain View, CA, USA
In-Office
Entry level
Information Technology • Software
The Role
Build and scale pipelines that generate, grade, verify, and QA reinforcement-learning environments and agentic coding tasks. Create realistic coding worlds around production codebases, automate task creation into the thousands, develop throughput-enhancing infrastructure, and manage the full task lifecycle. Analyze failures, prevent reward hacking and grader loopholes, run workloads on GCP, and use coding agents to accelerate environment development and validation.
Summary Generated by Built In
About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.

Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.

 
About the Role

You will own coding environments end to end. You choose what world to build, design the tasks inside it, build the grading, run frontier models against it, and harden it until the only way to pass is to actually do the work.

This is a delivery role. We care most about whether you have done this before and can point to what came out of it. If you have built agentic coding environments or tasks at real volume and can tell us how many and how hard they were, we want to talk.

 
What You'll Do

Build high-fidelity coding worlds around real codebases, with the conventions, dependencies, tooling, and accumulated mess that real software has.

Choose which environments are worth building. A strong coding environment hits several marks:

  • Targets work where frontier models measurably struggle

  • Exercises real engineering, meaning navigation, diagnosis, sequencing, and design, and not just writing a function

  • Rests on a codebase with enough history and structure that shortcuts do not survive

  • Has a clear pass condition that a reviewer would agree with

  • Produces many varied tasks from a single world rather than one

Design tasks across the full lifecycle. Prompt, environment, grader, running frontier models, failure analysis, and iteration, until each task is rigorous, fair, and hard to game.

Build grading and sandboxed execution that is deterministic and cannot be gamed. Assume the model will try to pass without doing the work, and close the door before it finds it.

Remove whatever is slowing the team down. Build the internal tooling that makes everyone around you faster.

Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.

What We're Looking For

A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce.

Experience scaling that output through automation rather than through more people doing more manual work.

Strong software engineering fundamentals and fluency in several languages that holds up in production code.

Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it.

An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that.

A clear sense of what frontier coding agents can and cannot do, and where they cut corners.

Ownership. You build, debug, and ship without much supervision.

You May Be a Good Fit If You Also
  • Have worked on RL training systems, post-training, verifiers, or tool-use harnesses

  • Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure

  • Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment

  • Have contributed to a public agentic benchmark such as Terminal-Bench

  • Have open-source work that other people depend on

What We Offer
  • Location: Mountain View, CA (Onsite)

  • Base Salary: $250,000 – $300,000 USD / year

  • Additional Comp: 25% performance-based bonus + equity

Benefits & Perks

  • Health, dental, and vision coverage

  • 401(k)

  • Daily onsite lunch provided

  • Visa sponsorship and relocation support available

  • Direct impact on how the industry trains and evaluates agents

We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.

 

Skills Required

  • Built and shipped pipelines producing RL environments or agentic tasks
  • Scaled task or environment creation into the hundreds or thousands through automation
  • Strong software engineering fundamentals
  • Fluency in Python
  • Production software experience with large codebases, build systems, testing, DevOps, SRE, diagnosis, and root-cause analysis
  • Ability to build pipelines, automation, grading and verification systems, and sandboxed execution infrastructure
  • Experience running workloads at scale on GCP
  • Hands-on experience with RL training systems, post-training, verifiers, or tool-use harnesses
  • Background in developer tooling, CI/CD sandboxes, or code-execution infrastructure
  • Experience reviewing task-creation processes for benchmarks such as Terminal Bench 3.0
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, California
13 Employees

What We Do

Bespoke Labs is a venture funded startup creating AI tools for data curation and post-training LLMs. (We are hiring!)

Similar Jobs

AfterQuery Logo AfterQuery

Software Engineer

Artificial Intelligence • Big Data
In-Office
San Francisco, CA, USA
200 Employees
200K-200K Annually

Mercor Logo Mercor

Full-stack Engineer

Artificial Intelligence • Software
In-Office
2 Locations
2217 Employees

Labelbox Logo Labelbox

Forward Deployed Engineer, RL Environments

Artificial Intelligence • Information Technology • Machine Learning
In-Office or Remote
7 Locations
115 Employees
140K-200K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account