AI Evaluators: Assessing A Shopping Assistant

Posted 14 Hours Ago
Be an Early Applicant
Hiring Remotely in United States
Remote
50-50 Hourly
Entry level
Artificial Intelligence • Information Technology • Software • Automation
The Role
Evaluate conversations between users and an AI shopping assistant, identify logical errors, inaccuracies, and poor recommendations, and create structured rubrics and verifiers for judging response quality. This remote engagement requires at least 20 hours per week and focuses on improving model performance and e-commerce response quality.
Summary Generated by Built In
What We're Researching

We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.

How It Works

You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.

Who This Is For

This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.

What You'll Do
  • Review real user interaction traces with an AI shopping assistant

  • Identify logical failures, inaccuracies, or poor recommendations in the text

  • Create structured rubrics and verifiers to judge response quality

  • Commit to 20+ hours per week of evaluation work on our internal platform

Who Should Apply
  • Experience in data evaluation, quality assurance, or AI training

  • Strong analytical skills with the ability to spot subtle errors in text

  • Familiarity with e-commerce search and digital shopping experiences

  • Ability to commit to a sustained workload of 20+ hours per week

Compensation

$50 per hour

 
Ready to participate?

Start your paid interview now

 
About Terac

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

 

Learn more at terac.com or on YouTube at @jointerac.

Skills Required

  • Experience in data evaluation, quality assurance, or AI training
  • Strong analytical skills and ability to identify subtle errors in text
  • Familiarity with e-commerce search and digital shopping experiences
  • Ability to commit to a sustained workload of 20 or more hours per week
  • Experience with prompt engineering, complex data annotation, or software testing
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
4 Employees
Year Founded: 2025

What We Do

Democratizing the future of company automation with AI

Similar Jobs

Samsara Logo Samsara

Customer Success Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
98K-148K Annually

Samsara Logo Samsara

Customer Success Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
88K-118K Annually

DigitalOcean Logo DigitalOcean

Technical Account Manager

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
In-Office or Remote
Austin, TX, USA
1400 Employees
141K-200K Annually

DraftKings Logo DraftKings

Corporate Counsel

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Remote or Hybrid
United States
6400 Employees
161K-201K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account