Research Scientist

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
In-Office
Junior
Artificial Intelligence
Scale AI for reinforcement learning environments
The Role
Own the measurement and improvement of model learning from tasks. Collaborate with frontier labs to design post-training recipes and data-quality techniques, run in-house evaluation systems, develop datasets and data products, assess acquired data, and build scalable ingestion and evaluation systems. The role requires production reinforcement learning experience at a frontier lab and end-to-end LLM post-training expertise.
Summary Generated by Built In

About idler

idler is a frontier data research lab. We build the evals and environments that the world's leading frontier labs use to measure and train their models.

After raising a $9m seed round led by Paradigm, we spent the last year developing coding evals for top coding models you know and love. At the same time, we've expanded into other domains besides coding: RSI & Auto-Research, Law, Enterprise Business, Cybersecurity, and others. Now, we are facing more lab demand for our data than we can serve, and are rapidly scaling the team to grow the business.

Our approach to creating training data scales using technology, and all of our data products are built on a unified self-reinforcing platform that learns through experience.

You would be joining a close-knit team that has reached product market fit, and your work would directly help to multiply our revenue.

You can see some of our work here: https://idler.ai/collections

About the role

As a Research Scientist at idler, you'll own measuring and improving how models learn from our tasks. The job is to maximize learning signal we produce per unit time. You'll draw on your own experience and collaborate with researchers at the frontier to validate our data quality, identify where improvements are needed, and create new datasets. To succeed, you'll need to have extensive experience doing this work in production at a frontier lab.

Examples of what you’ll do

  • Work with our customers — researchers at frontier labs — to design novel post-training recipes and data quality measurement techniques

  • Design and run our in-house post-training stack to measure model lift on our tasks

  • Develop new data products based on datasets and experts available to us

  • Identify opportunities to take advantage of self-reinforcing exponential feedback loops

  • Create agents to analyze thousands of environments and millions of trajectories

  • Help curate and specify task distributions for new corpora

  • Work with procurement to ensure external data we acquire is suitable for refinement

  • Create scaleable systems for ingesting & evaluating data we are considering buying

  • Develop new techniques for mining data for signal

What we’re looking for

  • 1+ years of experience doing RL in production at a frontier lab

  • Track record of post-training an LLM end to end

  • Desire to drive the research roadmap and implementation on a fast-moving team

  • Deep curiosity about how machines learn from data and how to extract the maximum learning signal from our tasks

Tech stack

Typescript, React, NodeJS, Postgres, Redis, Vercel, Cursor/Claude Code/Codex, Tinker, Modal, AWS, Daytona, GRPO

Details

  • In-person in San Francisco

  • Competitive salary + meaningful equity

  • Free meals in office

  • Healthcare, 401(k), 15 days of PTO per year

  • Relocation assistance

  • Small, ambitious team

This is an in-person role in San Francisco. We're a tight-knit founding team and we play to win. Join us if you like to win too.

Skills Required

  • 1+ years of experience doing reinforcement learning in production at a frontier lab
  • Track record of post-training a large language model end to end
  • Ability and desire to drive the research roadmap and implementation on a fast-moving team
  • Deep curiosity about how machines learn from data and how to extract maximum learning signal from tasks
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
6 Employees
Year Founded: 2025

What We Do

Idler builds reinforcement learning environments that teach AI models to code at expert human levels. We create training environments based on real-world coding scenarios that prepare models for the complex challenges they'll face in production.

Why Work With Us

You'll be joining a team that's expanding quickly, get direct access to AI researchers at frontier labs, and be put in a position to grow as fast as you can handle.

Gallery

Gallery

Idler Offices

OnSite Workspace

All employees work in person out of our office in the Dogpatch neighborhood of San Francisco.

Typical time on-site: None
HQSan Francisco, California
Dogpatch

Similar Jobs

Idler Logo Idler

Create Your Own Role

Artificial Intelligence
In-Office
San Francisco, CA, USA
6 Employees

Idler Logo Idler

Design Engineer

Artificial Intelligence
In-Office
San Francisco, CA, USA
6 Employees
125K-300K Annually

Idler Logo Idler

Forward Deployed Engineer

Artificial Intelligence
In-Office
San Francisco, CA, USA
6 Employees

Idler Logo Idler

Special Projects Operator

Artificial Intelligence
In-Office
San Francisco, CA, USA
6 Employees
140K-250K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account