About the company
Our client is a fast-growing technology company.
The role
Raydar is recruiting for this role on behalf of our client. Own the quality measurement framework that determines whether models and software features are ready to release. You will build test suites, reporting tools and human review rubrics, and partner with research, product and business colleagues so that quality is measurable and trusted.
What you'll do
- Build and maintain evaluation suites that act as a gate for each release, covering capability, behavior and regressions.
- Design human-rated review rubrics that catch issues automated checks miss.
- Develop dashboards and tooling that speed up experiment cycles and support clear decisions.
- Define and uphold the quality standard for what is ready to ship, including under deadline pressure.
- Work with researchers on measuring what matters, with product engineers on instrumenting real usage, and with external partners on explaining improvement in concrete terms.
- Audit existing evaluations early on, document what is useful, noisy or missing, and close the most significant gaps.
- Stand up new evaluation areas in the first months and help shape the research roadmap.
Requirements
What we're looking for
- 3 to 6 years of experience building evaluation frameworks for AI systems whose outputs vary from run to run.
- Strong engineering skills in tooling, dashboards, instrumentation and fast experiment loops.
- Ability to design rubrics for human reviewers.
- Confidence to challenge a gamed metric while keeping the team aligned.
- High agency and a strong sense of urgency.
- Track record of owning a quality bar and protecting it under product pressure.
- Hands-on evaluation experience in a real-world setting, such as a data or evaluation company or an advanced AI research group.
Bonus points
- Experience evaluating agentic, on-device or tool-using systems.
- Evaluation work at an AI research lab or a data and evaluation company.
- Exposure to consumer AI or device-maker quality assurance.
- Early-hire or fast-growing company experience.
Benefits
Compensation and benefits
- Base salary: USD 200,000 to 250,000 per year
- Equity
- Relocation support
Location and work model
- San Francisco, CA, United States
- On-site
- Full-time
Skills Required
- 3 to 6 years of experience building evaluation frameworks for AI systems with variable outputs
- Strong engineering skills in tooling, dashboards, instrumentation, and fast experiment loops
- Ability to design rubrics for human reviewers
- Ability to challenge gamed metrics while maintaining team alignment
- High agency and a strong sense of urgency
- Track record of owning and protecting a quality standard under product pressure
- Hands-on evaluation experience in a real-world setting, such as a data or evaluation company or advanced AI research group
- Experience evaluating agentic, on-device, or tool-using systems
- Evaluation work at an AI research lab or data and evaluation company
- Exposure to consumer AI or device-maker quality assurance
- Early-hire or fast-growing company experience
What We Do
Raydar is a talent acquisition and business consulting firm that connects world-class and emerging talent with growing organizations. It supports companies through team development, strategic hiring, and customized growth solutions, helping clients recruit roles such as engineers, product managers, executives, legal counsel, and quantitative traders. Raydar focuses on understanding each organization’s needs, culture, and long-term goals to build high-impact teams.






