Research Engineer

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
In-Office
200K-250K Annually
Mid level
Professional Services • Consulting
The Role
Own the quality measurement framework for AI models and software features. Build evaluation suites, human review rubrics, dashboards, reporting tools, and instrumentation that measure capability, behavior, regressions, and real-world usage. Establish release quality standards, audit existing evaluations, close measurement gaps, and collaborate with research, product, business, and external partners. The role requires protecting quality standards under product deadlines and shaping future evaluation and research priorities.
Summary Generated by Built In

About the company

Our client is a fast-growing technology company.

The role

Raydar is recruiting for this role on behalf of our client. Own the quality measurement framework that determines whether models and software features are ready to release. You will build test suites, reporting tools and human review rubrics, and partner with research, product and business colleagues so that quality is measurable and trusted.

What you'll do

- Build and maintain evaluation suites that act as a gate for each release, covering capability, behavior and regressions.

- Design human-rated review rubrics that catch issues automated checks miss.

- Develop dashboards and tooling that speed up experiment cycles and support clear decisions.

- Define and uphold the quality standard for what is ready to ship, including under deadline pressure.

- Work with researchers on measuring what matters, with product engineers on instrumenting real usage, and with external partners on explaining improvement in concrete terms.

- Audit existing evaluations early on, document what is useful, noisy or missing, and close the most significant gaps.

- Stand up new evaluation areas in the first months and help shape the research roadmap.


Requirements

What we're looking for

- 3 to 6 years of experience building evaluation frameworks for AI systems whose outputs vary from run to run.

- Strong engineering skills in tooling, dashboards, instrumentation and fast experiment loops.

- Ability to design rubrics for human reviewers.

- Confidence to challenge a gamed metric while keeping the team aligned.

- High agency and a strong sense of urgency.

- Track record of owning a quality bar and protecting it under product pressure.

- Hands-on evaluation experience in a real-world setting, such as a data or evaluation company or an advanced AI research group.

Bonus points

- Experience evaluating agentic, on-device or tool-using systems.

- Evaluation work at an AI research lab or a data and evaluation company.

- Exposure to consumer AI or device-maker quality assurance.

- Early-hire or fast-growing company experience.


Benefits

Compensation and benefits

- Base salary: USD 200,000 to 250,000 per year

- Equity

- Relocation support

Location and work model

- San Francisco, CA, United States

- On-site

- Full-time

Skills Required

  • 3 to 6 years of experience building evaluation frameworks for AI systems with variable outputs
  • Strong engineering skills in tooling, dashboards, instrumentation, and fast experiment loops
  • Ability to design rubrics for human reviewers
  • Ability to challenge gamed metrics while maintaining team alignment
  • High agency and a strong sense of urgency
  • Track record of owning and protecting a quality standard under product pressure
  • Hands-on evaluation experience in a real-world setting, such as a data or evaluation company or advanced AI research group
  • Experience evaluating agentic, on-device, or tool-using systems
  • Evaluation work at an AI research lab or data and evaluation company
  • Exposure to consumer AI or device-maker quality assurance
  • Early-hire or fast-growing company experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
28 Employees
Year Founded: 2021

What We Do

Raydar is a talent acquisition and business consulting firm that connects world-class and emerging talent with growing organizations. It supports companies through team development, strategic hiring, and customized growth solutions, helping clients recruit roles such as engineers, product managers, executives, legal counsel, and quantitative traders. Raydar focuses on understanding each organization’s needs, culture, and long-term goals to build high-impact teams.

Similar Jobs

Medra Logo Medra

Research Engineer, Post-training

Artificial Intelligence • Robotics • Software
In-Office
San Francisco, CA, USA
42 Employees

Exa (exa.ai) Logo Exa (exa.ai)

Research Engineer, Index Intelligence

Artificial Intelligence • Software
In-Office
San Francisco, CA, USA
86 Employees
180K-350K Annually

Exa (exa.ai) Logo Exa (exa.ai)

Research Engineer, Content Understanding

Artificial Intelligence • Software
In-Office
San Francisco, CA, USA
86 Employees
180K-350K Annually

Raydar Logo Raydar

Senior Research Engineer, Robotics

Professional Services • Consulting
In-Office
San Francisco, CA, USA
28 Employees
260K-300K Annually

Similar Companies Hiring

Fora Thumbnail
Agency • On-Demand • Professional Services • Sales • Software • Travel • Hospitality
New York, NY
250 Employees
Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
150 Employees
Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account