Research Engineer

Reposted 5 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Mid level
Artificial Intelligence • Software
The Role
Build scalable systems to aggregate, index, and analyze large-scale agent interaction data; develop agent-based evaluation systems; design post-training and optimization workflows; build internal tools and infrastructure for experimentation, analysis, and training; ship research into production and collaborate cross-functionally to improve agent behavior.
Summary Generated by Built In
The Role:

We are looking for Research Engineers to build AI systems that use agent interaction data to understand how agents behave, evaluate them at scale, and improve them through learning and feedback.

Your research will not live on a whiteboard. You’ll work directly with real-world agent data, apply frontier methods in production, and see your work ship into the product. By making agent behavior measurable and debuggable, your systems will support teams deploying agents across finance, legal, operations, and other high-stakes workflows. You will own projects end-to-end, with significant autonomy, and work closely with the team to build self-improving agent systems.

What You'll Do:
  • Build AI systems to aggregate, index, and analyze large-scale long-running agent interaction data in order to extract meaningful signals

  • Design and implement post-training and optimization workflows to improve agents, both internally and for customers

  • Build agent platform infrastructure, including orchestration, runtimes, and developer tools that help teams define, test, deploy, and iterate on complex agent workflows

  • Build internal tools and infrastructure that support rapid experimentation, analysis, and training

  • Work closely with product to integrate agents into customer-facing workflows

  • Collaborate with external companies and research partners on frontier AI research

What We're Looking For

Every hire clears three bars, no exceptions:

  • Agency. You are intellectually curious, self-directed, and stay up to date with the latest research, blogs, trends, and ideas.

  • Depth of thought. You can reason clearly about abstract systems, and ideally have experience working on agents, RL, or the infrastructure that supports them.

  • Ownership. You own outcomes, not just tasks. You use freedom to experiment responsibly, make business-driven decisions, and focus first on work that moves the company forward.

More specifically, you should bring strength in at least one of the following areas:

  • Data quality, evaluation, benchmarking, and hands-on work with messy production data

  • Agent systems built or evaluated in real-world or production settings

  • Reinforcement learning, post-training, agents, or machine learning fundamentals

  • Infrastructure and systems work across training, data pipelines, evaluation, or model serving

  • Translating research into product while balancing customer constraints, technical tradeoffs, and business impact

  • Turning ambiguous problems into clear, well-designed plans

Skills Required

  • Care about data quality, evaluation, and benchmarking; comfortable working hands-on with messy data.
  • Experience building agent systems and working with them in real-world or production settings.
  • Strong background in reinforcement learning, agents, or machine learning fundamentals.
  • Comfortable across infrastructure and systems, spanning training, data pipelines, and model serving.
  • Able to translate research into product and balance real-world customer constraints and tradeoffs across teams.
  • Ability to turn ambiguous problems into clear, well-designed plans and own projects end-to-end.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
20 Employees
Year Founded: 2025

What We Do

Judgment Labs builds agent behavior monitoring (ABM) infrastructure. Judgment provides a toolkit to track and judge agent behavior in online and offline setups, enabling you to convert high-signal interaction data from production/test environments into more reliable agents.

Similar Jobs

HRL Laboratories Logo HRL Laboratories

Agentic AI & Graph Machine Learning Research Engineer

Artificial Intelligence • Hardware • Software • Nanotechnology • Semiconductor • Quantum Computing • Defense
Hybrid
Calabasas, CA, USA
850 Employees
128K-160K Annually

CrowdStrike Logo CrowdStrike

Engineer II, Advanced Research (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
37 Locations
11000 Employees
100K-145K Annually

CoreWeave Logo CoreWeave

Staff Applied Research Engineer

Cloud • Information Technology • Machine Learning
In-Office
2 Locations
1450 Employees
207K-275K Annually

CoreWeave Logo CoreWeave

Senior Applied Research Engineer

Cloud • Information Technology • Machine Learning
In-Office
2 Locations
1450 Employees
182K-242K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account