AI Test Engineer - VLabs

Posted 5 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Mobile • Consulting
The Role
Design and execute AI evaluation and test frameworks for LLMs, OCR, VLM, and agentic workflows. Build reusable Python-based evaluation pipelines, drift detection, ground-truth datasets, and CI/CD-integrated tests. Validate outputs, detect hallucinations, perform API/SQL checks, and enable release gating while ensuring Responsible AI and stakeholder reporting.
Summary Generated by Built In

About Vialto Labs (VLabs) 

Vialto Labs (VLabs) is responsible for redesigning how work is delivered in the tax and immigration service lines, as well as driving operational efficiency across Vialto’s functional areas using AI. The team builds and deploys novel AI-enabled solutions that directly improve productivity and increase delivery quality for our clients. VLabs is accountable for rapidly turning innovative experiments into production-ready deliverables at scale and embedding them into day-to-day operations. This team focuses on the highest-impact workflows, creating standardized, repeatable capabilities that can be deployed globally. Operating with a mandate for speed and measurable outcomes, VLabs works alongside service line, product, and platform leaders. 

About the Role 

AI Test Engineering is a hands-on role within VLabs Quality Engineering, responsible for validating the performance, reliability, and integrity of AI-enabled solutions in production environments. This role operates at the intersection of AI engineering and quality assurance, ensuring that outputs from LLMs, OCR pipelines, document classification models, and agentic workflows perform as expected at scale and meet defined business performance thresholds.  Working closely with the Programme Test Manager and partnering with engineering, product, and delivery teams, this role translates AI testing strategy into executable frameworks, evaluation pipelines, and reusable assets embedded into the delivery lifecycle. 

Success requires independent execution, strong technical depth, and the ability to proactively identify risks, patterns, and performance gaps while enabling rapid, production-grade deployment of AI capabilities. 

Key Responsibilities 

AI Evaluation & Test Design 

  • Translate AI testing strategy into executable test scenarios across LLM outputs, document classification, extraction accuracy, agent workflows, and edge cases 

  • Design adversarial and boundary test inputs to expose hallucination, misclassification, and failure modes 

  • Validate AI outputs for structure, consistency, accuracy, and production readiness against defined performance thresholds 

Evaluation Engineering & Automation 

  • Build reusable Python-based evaluation frameworks, including output validation, hallucination detection, and scoring mechanisms 

  • Develop parameterized test scripts reusable across features, models, and releases 

  • Implement AI-as-Judge frameworks, including prompt design, scoring logic, and calibration of evaluation reliability 

  • Embed evaluation frameworks into CI/CD pipelines to support continuous testing and deployment 

Drift Detection & Quality Monitoring 

  • Design and operate drift detection frameworks using fixed baseline datasets and scheduled re-evaluation 

  • Establish thresholds to distinguish acceptable variation from performance degradation 

  • Enable release gating by identifying regressions prior to production deployment 

Ground Truth & Data Quality 

  • Build and maintain ground truth datasets in partnership with subject matter experts 

  • Define standards for classification, extraction accuracy, and acceptable output characteristics 

  • Continuously update datasets to reflect evolving business requirements and use cases 

Workflow & Integration Testing 

  • Test end-to-end agentic workflows, validating data integrity, error propagation, and fallback behavior 

  • Perform API-level testing of AI pipeline endpoints using Python and Postman/Newman 

  • Validate data persistence and integrity across system layers using SQL 

  • Partner with engineering teams to ensure testability, observability, and system reliability 

Standardization & Scaling 

  • Define and scale standardized AI evaluation patterns and reusable quality frameworks across VLabs 

  • Contribute to enterprise AI quality standards and reference architectures 

Governance & Responsible AI 

  • Ensure adherence to Responsible AI, data privacy, and governance requirements 

  • Support auditability, traceability, and transparency of AI outputs and evaluation processes 

Stakeholder Enablement 

  • Translate evaluation results into actionable insights for engineering, product, and business stakeholders 

  • Support decision-making on model readiness, release risk, and performance trade-offs 

  • Proactively identify risks, patterns, and systemic issues and escalate appropriately 

Qualifications & Experience 

Professional Experience 

  • 7+ years in software testing, including 2–3 years focused on AI/ML-enabled systems in production environments 

  • Proven experience designing and executing AI evaluation frameworks and quality strategies 

  • Strong track record building ground truth datasets, drift detection systems, and scalable evaluation pipelines 

  • Experience testing multi-step agentic workflows and AI-driven automation systems 

  • Experience operating in fast-paced, iterative delivery environments 

  • Background in regulated or compliance-driven environments preferred 

Technical Expertise 

  • Advanced Python programming for evaluation frameworks, batch processing, and data analysis 

  • Experience with LLM evaluation tools such as deepeval, RAGAS, promptfoo, or similar 

  • Strong capabilities in: 

  • AI output validation, hallucination detection, and grounding checks 

  • Drift detection frameworks and statistical evaluation methods 

  • OCR, VLM, and document AI testing (classification, extraction, edge cases) 

  • API testing using Python (requests/httpx) and Postman/Newman 

  • SQL for data validation and pipeline integrity checks 

  • Familiarity with LangChain, LlamaIndex, or similar frameworks 

  • Experience with cloud AI platforms such as Azure AI Foundry or AWS Bedrock preferred 

Operating Capabilities 

  • Ability to operate independently in fast-moving, ambiguous environments 

  • Strong analytical mindset with attention to detail and quality rigor 

  • Ability to balance speed and rigor in AI evaluation and delivery cycles 

  • Proactive communicator who identifies risks and drives resolution 

  • Ability to translate technical findings into business-relevant insights 

Education 

  • Bachelor’s degree required

  • Advanced degree in Computer Science, Data Science, or related field preferred 

Additional Information:

  • This job is based in our Bangalore office with the possibility of hybrid work mode

  • We are an equal opportunity employer that does not discriminate on the basis of any legally protected status

  • Please note, AI is used as part of the application process

Skills Required

  • 7+ years in software testing, including 2-3 years focused on AI/ML-enabled systems in production
  • Proven experience designing and executing AI evaluation frameworks and quality strategies
  • Track record building ground truth datasets, drift detection systems, and scalable evaluation pipelines
  • Experience testing multi-step agentic workflows and AI-driven automation systems
  • Background in regulated or compliance-driven environments
  • Advanced Python programming for evaluation frameworks, batch processing, and data analysis
  • Experience with LLM evaluation tools such as deepeval, RAGAS, promptfoo, or similar
  • AI output validation, hallucination detection, and grounding checks
  • Drift detection frameworks and statistical evaluation methods
  • OCR, VLM, and document AI testing (classification, extraction, edge cases)
  • API testing using Python (requests/httpx) and Postman/Newman
  • SQL for data validation and pipeline integrity checks
  • Familiarity with LangChain, LlamaIndex, or similar frameworks
  • Experience with cloud AI platforms such as Azure AI Foundry or AWS Bedrock
  • Bachelor's degree
  • Advanced degree in CS, Data Science, or related field
  • Ability to operate independently in fast-moving, ambiguous environments; strong analytical and communication skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
5,845 Employees

What We Do

Vialto Partners is focused on transforming the global mobility ecosystem to support multinational organizations and their employees. The organization sparks growth and creates meaningful impact for businesses and individuals in more than 150 countries, supported by a team of over 6,000 dedicated immigration, tax, human resources and technology professionals. ​ Vialto Partners offers clients:​ Tax solutions: Strategic tax planning, preparation and cross-border compliance services, including tax equalization and employment structures.​ Immigration services: Immigration-related consulting, processing and compliance services informed by knowledge of government regulations and rapidly changing immigration laws.​ Compensation and rewards: Integrated, end-to-end payroll and compensation management and reporting to ensure holistic global reporting, local compliance and overall cost-effectiveness.​ Dynamic work strategy and services: Design and implementation of real-time mobility strategies and services—from mobility managed services to remote work and business travel solutions.​ Vialto Partners is reimagining the future of global mobility, driving the mobility ecosystem forward to create a more connected, integrated and efficient supply chain to meet and exceed the evolving challenges and complexities of global workforce management. #vialtopartners​

Similar Jobs

CSC Logo CSC

Accountant

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
8500 Employees

GitLab Logo GitLab

Rpa Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
2500 Employees

LogicMonitor Logo LogicMonitor

Senior Sales Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Remote or Hybrid
India
1100 Employees

eClinical Solutions Logo eClinical Solutions

Senior Data Engineer

Cloud • Healthtech • Professional Services • Software • Pharmaceutical
Easy Apply
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
400 Employees

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
ARB Interactive Thumbnail
Gaming • Mobile • Software
Miami, Florida
190 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account