Senior AI QA Engineer

Posted Yesterday
Be an Early Applicant
4 Locations
In-Office
Senior level
Information Technology • Software • Analytics
The Role
Lead QA for enterprise LLM, RAG, and AI agent applications: design benchmarks, build automated evaluation and observability pipelines, validate retrieval and agent workflows, implement Python/Pytest automation, create Playwright UI tests, integrate quality gates into CI/CD, perform performance validation and root cause analysis to ensure production readiness.
Summary Generated by Built In

Job Position : Senior AI Quality Assurance (QA) Engineer
Location : Bengaluru
Experience : 7+ Years 
Required Skills:

We are looking for an experienced Senior AI Quality Assurance (QA) Engineer to ensure the quality, reliability, and production readiness of enterprise-grade AI applications.

This role extends beyond traditional software testing and focuses on the evaluation and quality engineering of LLM applications, RAG systems, and AI agents. You will own automated testing, AI evaluations, benchmark creation, observability, prompt regression testing, and end-to-end validation of AI workflows. Working closely with AI Engineers, Product Managers, and Platform teams, you will establish measurable quality standards and ensure every release meets enterprise-grade expectations for accuracy, reliability, performance, and scalability.

Key Responsibilities:

AI Evaluation & Benchmarking

Design benchmark (golden) datasets and build automated evaluation pipelines for LLM applications. Define quality gates and continuously evaluate prompts, models, retrieval pipelines, and agent behaviour using metrics such as hallucination rate, tool selection accuracy, execution accuracy, precision, recall, latency, and cost.

RAG & Agent Quality Validation

Validate Retrieval-Augmented Generation (RAG) pipelines and AI agent workflows by testing retrieval quality, context relevance, tool invocation, reasoning flow, memory, and end-to-end task completion. Design evaluation scenarios covering ambiguous queries, multi-turn conversations, retrieval failures, and edge cases.

Python Automation & API Testing

Develop and maintain scalable automation frameworks using Python and Pytest for unit testing, integration testing, API testing, regression testing, and end-to-end validation. Build reusable test utilities and integrate automated quality checks into CI/CD pipelines.

Frontend Automation

Develop automated UI test suites using Playwright to validate AI-powered user journeys, conversational interfaces, workflow execution, and end-to-end application behaviour across releases.

Observability & Root Cause Analysis

Use OpenTelemetry, tracing platforms, and AI observability tools to analyse execution traces, latency, model responses, API calls, and workflow behaviour. Perform root cause analysis to identify regressions, hallucinations, bottlenecks, and production issues.

Performance & Enterprise Readiness

Validate AI application performance by monitoring latency, throughput, reliability, and scalability. Ensure production readiness through regression testing, API validation, workflow testing, and enterprise quality standards.

Required Skills:

7+ years of experience in Software QA, Test Automation, or AI Quality Engineering.

Strong Python programming skills.

Hands-on experience with Pytest for:

Unit testing

Integration testing

API testing

Regression testing

Experience testing REST APIs and backend services.

Experience evaluating LLM-powered applications.

Hands-on experience evaluating Retrieval-Augmented Generation (RAG) systems.

Understanding of AI agent evaluation methodologies.

Experience measuring AI quality using metrics such as:

Hallucination Rate

Tool Selection Accuracy

Execution Accuracy

Precision / Recall

Latency (P50/P95/P99)

Token Usage

Experience with Playwright or similar frontend automation frameworks.

Experience with OpenTele metry, tracing, or observability platforms.

Strong debugging and root cause analysis skills.

Experience integrating automated tests into CI/CD pipelines.

Nice to Have

Experience with Lang Smith, Lang fuse, MLflow, Arize Phoenix, or similar AI observability platforms.

Experience evaluating multi-agent systems and orchestration frameworks such as LangGraph, CrewAI, Google ADK or AutoGen.

Experience with vector databases such as Pinecone, Milvus, Weaviate, pgvector, or Vertex AI Vector Search.

Exposure to OpenAI, Anthropic, Gemini, or Azure OpenAI.

Experience with performance and load testing tools.

Prior experience in Banking, Financial Services, or Insurance (BFSI).

Skills Required

  • 7+ years of experience in Software QA, Test Automation, or AI Quality Engineering
  • Strong Python programming skills
  • Hands-on experience with Pytest (unit, integration, API, regression testing)
  • Experience testing REST APIs and backend services
  • Experience evaluating LLM-powered applications
  • Hands-on experience evaluating Retrieval-Augmented Generation (RAG) systems
  • Understanding of AI agent evaluation methodologies
  • Experience measuring AI quality using metrics (hallucination rate, tool selection accuracy, execution accuracy, precision/recall, latency, token usage)
  • Experience with Playwright or similar frontend automation frameworks
  • Experience with OpenTelemetry, tracing, or observability platforms
  • Strong debugging and root cause analysis skills
  • Experience integrating automated tests into CI/CD pipelines
  • Experience with LangSmith, Langfuse, MLflow, Arize Phoenix, or similar AI observability platforms
  • Experience evaluating multi-agent systems and orchestration frameworks (LangGraph, CrewAI, Google ADK, AutoGen)
  • Experience with vector databases (Pinecone, Milvus, Weaviate, pgvector, Vertex AI Vector Search)
  • Exposure to OpenAI, Anthropic, Gemini, or Azure OpenAI
  • Experience with performance and load testing tools
  • Prior experience in Banking, Financial Services, or Insurance (BFSI)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Wixom, MI
28 Employees
Year Founded: 2003

What We Do

DATAMAXIS takes pride in delivering a wide range of business IT modernization, data analytics, and technology management services. With command of the cutting-edge developments in these fields, our team and consultants are ready to provide you a robust technology modernization experience that results in a big boost in performance capability and operational efficiency.

Similar Jobs

CrowdStrike Logo CrowdStrike

Infrastructure Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees

Capco Logo Capco

RWA

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Capco Logo Capco

Liquidity Reporting

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Mondelēz International Logo Mondelēz International

Organization Capability Specialist

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
India
90000 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account