The Role
Lead AI-focused QA efforts by evaluating LLM and conversational AI systems, designing benchmark datasets, implementing automated tests with Python and Pytest, validating REST APIs and distributed services, verifying NL2SQL outputs and SQL correctness, using AI observability platforms, and performing root-cause analysis across complex AI workflows.
Summary Generated by Built In
Job Position: Senior AI Quality Assurance (QA) Engineer
Experience: 7+ Years
Location: Bengaluru
Required Skills
- 7+ years of experience in Software Quality Assurance, Test Automation, or AI Quality Engineering.
- Hands-on experience evaluating LLM-powered applications, conversational AI systems, NL2SQL solutions, or similar AI workflows.
- Strong understanding of AI benchmarking methodologies and experience creating benchmark (golden) datasets for regression testing.
- Experience with AI evaluation and observability platforms such as Lang Smith, MLflow, Arize Phoenix, or equivalent.
- Strong understanding of AI evaluation metrics, including Precision, Recall, Precision@K, Recall@K, F1 Score, Exact Match (EM), Mean Reciprocal Rank (MRR), Execution Accuracy, latency (P50/P95/P99), token usage, and cost analysis.
- Strong proficiency in Python with hands-on experience building automation using Pytest, including unit testing, integration testing, API testing, exception handling, mocking, fixtures, and parameterized testing.
- Experience testing REST APIs and distributed backend services.
- Strong SQL skills with the ability to validate generated queries, execution results, schema alignment, and business logic.
- Strong analytical and debugging skills with the ability to perform root cause analysis across complex AI systems.
Skills Required
- 7+ years of experience in Software Quality Assurance, Test Automation, or AI Quality Engineering.
- Hands-on experience evaluating LLM-powered applications, conversational AI systems, NL2SQL solutions, or similar AI workflows.
- Strong understanding of AI benchmarking methodologies and experience creating benchmark (golden) datasets for regression testing.
- Experience with AI evaluation and observability platforms such as LangSmith, MLflow, Arize Phoenix, or equivalent.
- Strong understanding of AI evaluation metrics including Precision, Recall, Precision@K, Recall@K, F1 Score, Exact Match, MRR, Execution Accuracy, latency (P50/P95/P99), token usage, and cost analysis.
- Strong proficiency in Python with hands-on experience building automation using Pytest, including unit, integration, API testing, exception handling, mocking, fixtures, and parameterized testing.
- Experience testing REST APIs and distributed backend services.
- Strong SQL skills with ability to validate generated queries, execution results, schema alignment, and business logic.
- Strong analytical and debugging skills with ability to perform root cause analysis across complex AI systems.
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
DATAMAXIS takes pride in delivering a wide range of business IT modernization, data analytics, and technology management services. With command of the cutting-edge developments in these fields, our team and consultants are ready to provide you a robust technology modernization experience that results in a big boost in performance capability and operational efficiency.








