Head of Evaluations (Legal AI Benchmarking)

Reposted One Month Ago
4 Locations
Remote
Senior level
Artificial Intelligence • Legal Tech • Natural Language Processing • Software
The Role
Lead design and scaling of legal AI evaluation frameworks and benchmarks, build and maintain datasets, audit and score AI legal outputs, define metrics focused on reasoning and citation accuracy, and collaborate with engineering to translate legal errors into model fine-tuning feedback.
Summary Generated by Built In
Who are we?

At Newcode.ai, we're transforming how law firms and legal professionals harness AI for real-world impact. As part of our collaborative, high-growth team, you'll have the rare opportunity to work side-by-side with visionary founders at the bleeding edge of AI and legal innovation — shaping not just our product, but the future of legal work itself.

Note: We believe in being transparent about what it's like to work at Newcode. As a fast-growing startup, we're building and evolving every day. That means not every process, playbook, or framework is already in place, and priorities can shift quickly.

The people who thrive here are comfortable with ambiguity, take ownership, and don't wait for perfect direction. They are resourceful, proactive, and able to "figure it out"—solving problems, creating structure where needed, and helping build the company as they go. If you are good with this then, great! Keep reading to learn more.

Position Overview 

We are seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform. You will ensure our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate. 

Key Responsibilities 

  • Design Legal Benchmarks for: Contract Drafting, Information Extraction, Legal Research, and Contract Review 
  • Build, source and maintain relevant datasets 
  • Audit AI Output: Review and score complex AI-generated legal text, contract analyses, and statutory interpretations for accuracy and precision and lay out a strategy.  
  • Define Evaluation Metrics: Establish clear criteria for grading model performance, specifically focusing on logical reasoning, citation accuracy, and the model's ability to safely abstain from answering. 
  • Collaborate with Engineering: Partner directly with Engineering to translate legal errors into actionable technical feedback for model fine-tuning. 

Requirements
  • PhD or Masters in statistics, mathematics, machine learning or equivalent 
  • Analytical Skills: Proven ability to break down complex statutory frameworks and case law into structured, logical data points. 
  • Tech-Savviness: python, panda, numpy, jupiter notebooks and similar statistical models 

Visa Sponsorship:

At this time, Newcode is unable to provide visa sponsorship. Candidates must be authorized to work in the applicable country without employer sponsorship.

Skills Required

  • PhD or Masters in statistics, mathematics, machine learning or equivalent
  • Strong statistical background and analytical skills to structure complex statutory frameworks and case law
  • Experience designing legal benchmarks for contract drafting, information extraction, legal research, and contract review
  • Experience building, sourcing, and maintaining datasets for model evaluation
  • Technical proficiency with Python, pandas, numpy, and Jupyter Notebooks (statistical modeling workflows)
  • Ability to audit and score complex AI-generated legal text for factual and legal accuracy
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
13 Employees
Year Founded: 2021

What We Do

Newcode.ai is an AI-native platform for legal professionals that automates complex, multi-step legal workflows and document analysis.

Similar Jobs

Pfizer Logo Pfizer

Senior Manager, HTA, Value and Evidence (HV&E), Genitourinary Cancer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
139K-232K Annually

Pfizer Logo Pfizer

Staff Software Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees

Pfizer Logo Pfizer

Sustainability Senior Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
30 Locations
121990 Employees
112K-207K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Client Delivery Lead Top Accounts

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office or Remote
7 Locations
85422 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account