Full stack engineer

Posted 22 Days Ago
San Francisco, CA, USA
In-Office
Senior level
Artificial Intelligence • Software
The Role
Build end-to-end product and agent systems that ingest production agent traces, run large-scale investigations, verify agent changes via simulated envs and trajectory replay, design investigator and swarm UX, and build platform features (workspaces, roles, billing, SDKs) to close the improvement loop for AI agents.
Summary Generated by Built In
Full-Stack Engineer

San Francisco · On Site · Full Time

The Role

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:

  1. We ingest everything your agents do in production: traces, tool calls, decisions, outcomes

  2. Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals

  3. Teams close the loop, shipping agent improvements validated against real production evidence

You'll build the product experiences that make this loop legible, and you'll build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.

What You Will Accomplish
  • Judgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.

  • Verification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes.

  • Agent investigation interfaces: Design how engineers understand what their agents did and why. Long traces, tool calls, decisions, failures. What does debugging look like when the "program" is a reasoning loop? How do you make a thousand-step trajectory legible in minutes?

  • Swarm UX: A hundred parallel investigations is useless if engineers can't follow them. Design how humans watch a swarm work, redirect investigators that go down the wrong path, and consume findings without reading a hundred reports.

  • The improvement loop: Build the workflows that turn production trajectories into datasets, judges, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion.

  • The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.

  • Judgment everywhere agents are built: An SDK and terminal-first experience so Claude Code, Codex, and OpenCode sessions can summon Judgment as a subagent mid-development.

What You'll Bring
  • Experience building and scaling end-to-end production systems, from data layer to UI

  • Strong technical problem-solving skills, especially in fast-changing, ambiguous environments

  • A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship

  • Hands-on experience building with LLMs or agents, or the drive to get there fast

  • Comfort working directly with customers to understand their needs and solve real-world problems

  • Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences

Skills Required

  • Experience building and scaling end-to-end production systems from data layer to UI
  • Strong technical problem-solving skills in ambiguous, fast-changing environments
  • Builder mindset with high agency and ability to ship features independently
  • Hands-on experience building with LLMs or agents, or rapid ability to learn them
  • Comfort working directly with customers to gather requirements and iterate
  • Excellent communication skills across technical and non-technical audiences
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
20 Employees
Year Founded: 2025

What We Do

Judgment Labs builds agent behavior monitoring (ABM) infrastructure. Judgment provides a toolkit to track and judge agent behavior in online and offline setups, enabling you to convert high-signal interaction data from production/test environments into more reliable agents.

Similar Jobs

Navixus | Tech Mahindra Logo Navixus | Tech Mahindra

Full-stack Engineer

Artificial Intelligence • Natural Language Processing • Professional Services • Analytics • Consulting • Conversational AI • Generative AI
Hybrid
San Ramon, CA, USA
830 Employees

CrowdStrike Logo CrowdStrike

Full-stack Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
4 Locations
11000 Employees
100K-145K Annually

Stripe Logo Stripe

Full-stack Engineer

Payments • Software
In-Office
San Francisco, CA, USA
5360 Employees

SpaceX Logo SpaceX

Full-stack Engineer

Aerospace • Other
In-Office
Palo Alto, CA, USA
8879 Employees
135K-185K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account