QA/Test Engineer

Posted 2 Days Ago
Hiring Remotely in United States
Remote
60-90 Hourly
Junior
Artificial Intelligence • HR Tech • Professional Services • Software
The Role
Evaluate and test complex AI benchmark tasks by designing test cases, executing tasks, debugging environments, reviewing grading logic, identifying shortcuts, and developing repeatable QA processes while collaborating with researchers to maintain rigorous quality standards.
Summary Generated by Built In

This role is for one of our clients

Compensation: $60-$90 per hour

A leading AI research organization is developing the next generation of agentic evaluation benchmarks for advanced AI models. These complex, multi-step tasks need to be reliable, unambiguous, accurately graded, and resistant to shortcuts. We are seeking experienced QA/Test Engineers to help establish the quality standards and testing processes that ensure every benchmark task measures what it is intended to measure.

Each task may represent one to two days of expert development and can span multiple technical skills. As a QA/Test Engineer, you will deeply evaluate these tasks by executing them, testing edge cases, debugging environments, and identifying potential weaknesses before they reach production evaluation. You will work closely with researchers and task authors in an iterative feedback environment.

Work Arrangement: Fully Remote — United States
Commitment: Approximately 35 hours per week
Employment Type: Full-Time


RequirementsKey Responsibilities
  • Design Test Cases: Develop comprehensive test scenarios to verify that benchmark tasks function correctly, including edge cases, failure conditions, and unexpected inputs.
  • Review Task Quality: Thoroughly review task descriptions, requirements, expected outcomes, and reference solutions to identify ambiguity, inconsistencies, missing requirements, and potential quality issues.
  • Debug & Troubleshoot: Use Python and other development tools to investigate and resolve issues within task environments, validation logic, and automated checks.
  • Develop QA Processes: Create practical, repeatable testing procedures, quality checklists, and review frameworks that can be consistently applied across benchmark tasks.
  • Identify Evaluation Gaps: Examine grading logic and AI-agent execution results to detect loopholes, shortcuts, inconsistent scoring, or other factors that could compromise benchmark reliability.
  • Collaborate with Researchers: Work closely with researchers and task authors to communicate issues clearly, recommend improvements, and ensure fixes are properly validated.
  • Maintain Quality Standards: Help establish and enforce rigorous quality standards across complex technical evaluation datasets.
Core Qualifications
  • Education: MSc or PhD in a STEM discipline, or equivalent practical experience in a research-intensive or engineering-focused environment.
  • Experience: 1+ years of experience in QA, test engineering, software engineering, research engineering, or a comparable role involving significant ownership of quality.
  • Proven ability to design effective test cases, develop QA processes, and investigate complex technical systems end-to-end.
  • Working proficiency in Python and Git, with the ability to understand unfamiliar codebases, environments, and technical workflows.
  • Strong debugging and analytical skills with an ability to systematically isolate and resolve issues.
  • Exceptional attention to detail and a strong habit of maintaining clear, structured technical documentation.
  • Ability to identify subtle failures, inconsistencies, edge cases, and unintended behaviors that may be overlooked by others.
  • Previous experience with AI evaluation, AI training, model testing, benchmark development, or reviewing AI-generated outputs is highly desirable.
  • Comfortable working independently on ambiguous and open-ended technical problems.
  • Strong written communication skills and the ability to provide actionable feedback to technical stakeholders.
  • A quality-focused mindset with the creativity and persistence to uncover problems others may miss.
  • Ability to commit reliably to approximately 35 hours per week.
Ideal Candidate

The ideal candidate combines strong software testing and debugging capabilities with exceptional analytical judgment. You should enjoy breaking complex systems, investigating unexpected behavior, and asking whether a test genuinely proves what it claims to prove.

Experience working with AI systems, evaluation frameworks, automated grading, or complex technical benchmarks will be particularly valuable.

Skills Required

  • MSc or PhD in a STEM discipline or equivalent practical experience
  • 1+ years experience in QA, test engineering, software engineering, research engineering, or comparable quality-focused role
  • Proven ability to design effective test cases and develop QA processes
  • Working proficiency in Python
  • Working proficiency in Git
  • Strong debugging and analytical skills with ability to isolate and resolve issues
  • Exceptional attention to detail and clear, structured technical documentation habits
  • Ability to identify subtle failures, inconsistencies, edge cases, and unintended behaviors
  • Previous experience with AI evaluation, AI training, model testing, benchmark development, or reviewing AI-generated outputs
  • Comfortable working independently on ambiguous and open-ended technical problems
  • Strong written communication skills and ability to provide actionable feedback to technical stakeholders
  • Ability to commit reliably to approximately 35 hours per week
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2021

What We Do

Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

Similar Jobs

TIAG Logo TIAG

Test Engineer

Information Technology • Security
In-Office or Remote
20190, Reston, VA, USA
348 Employees
120K-160K Annually

Panum Group Logo Panum Group

Test Engineer

Information Technology
In-Office or Remote
McLean, VA, USA
160 Employees
95K-120K Annually

Groundswell Logo Groundswell

Test Engineer

Information Technology • Consulting
In-Office or Remote
50 Locations
360 Employees
83K-117K Annually

Weekday, Inc. Logo Weekday, Inc.

Test Engineer

Artificial Intelligence • HR Tech • Professional Services • Software
Remote
United States
60-90 Hourly

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account