Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:
We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals
Teams close the loop, shipping agent improvements validated against real production evidence
You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great. This is not a role where you implement specs handed down.
What You Will AccomplishInvestigation interfaces: Design how engineers understand what their systems did and why. Long traces, tool calls, decisions, failures. How do you make a complex sequence of events legible in minutes?
Verification: Build the platform for verifying system changes: hosted simulated environments, trajectory replay, and monitors for unintended behavior changes.
The improvement loop: Build the workflows that turn production data into datasets, evaluations, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion.
The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many workflows across many environments.
Experience building and scaling end-to-end production systems, from data layer to UI
Strong technical problem-solving skills, especially in fast-changing, ambiguous environments
A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship
Comfort working directly with customers to understand their needs and solve real-world problems
Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences
Skills Required
- Experience building and scaling end-to-end production systems from data layer to UI
- Strong technical problem-solving skills in ambiguous, fast-changing environments
- Builder mindset with high agency and ability to ship features independently
- Hands-on experience building with LLMs or agents, or rapid ability to learn them
- Comfort working directly with customers to gather requirements and iterate
- Excellent communication skills across technical and non-technical audiences
What We Do
Judgment Labs builds agent behavior monitoring (ABM) infrastructure. Judgment provides a toolkit to track and judge agent behavior in online and offline setups, enabling you to convert high-signal interaction data from production/test environments into more reliable agents.









