We are looking for someone who has validated AI systems in production - where outputs are non-deterministic, failure modes are ambiguous, and correctness is probabilistic.
You will be responsible for evaluating, hardening, and governing GenAI and agentic systems before and after they go live. This role sits at the intersection of engineering, QA, and AI safety, ensuring that what we build works consistently, reliably, and safely - at scale.
You will partner closely with AI engineers, architects, and clients to define what “correct” means in an AI world — and prove it.
ResponsibilitiesAI Quality Engineering & Evaluation
Define evaluation frameworks for LLMs, RAG pipelines, and multi-agent systems
Design test strategies for non-deterministic systems (semantic correctness, hallucination detection, consistency checks)
Build testing pipelines for LLM evaluation (offline + online evals)
Establish ground truth datasets, benchmarks, and scoring metrics (precision, recall, relevance, factuality)
Agent & Workflow Validation
Validate multi-agent orchestration flows.
Test failure modes: hallucination, incomplete reasoning.
Simulate edge cases and adversarial inputs.
Ensure robustness across multi-step workflows and chained reasoning tasks
RAG & Data Validation
Validate end-to-end RAG pipelines:
Chunking quality
Embedding correctness
Retrieval relevance
Re-ranking effectiveness
Detect and quantify RAG failure points (retrieval gaps, stale data, hallucinations)
Ensure data lineage and traceability in AI responses
Guardrails, Safety & Governance
Ensure compliance with enterprise AI governance and auditability requirements
Validate explainability and traceability of AI outputs
Automation & Tooling
- Automate prompt testing, regression testing, and response comparison
- Integrate AI validation into CI/CD pipelines
- Built or contributed to evaluation frameworks for LLM-based systems
- Tested RAG pipelines end-to-end and identified failure points
- Defined and executed non-deterministic test strategies
- Automated LLM evaluation or prompt regression pipelines
- Worked on agent-based or multi-step AI workflows
- Debugged incorrect or hallucinated model outputs using structured methods
- Established quality metrics where ground truth was unclear or evolving
Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. Explore Life at Zensar and join us to Grow. Own. Achieve. Learn. to be the best version of yourself.
We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Skills Required
- Built or contributed to evaluation frameworks for LLM-based systems
- Tested RAG pipelines end-to-end and identified failure points
- Defined and executed non-deterministic test strategies
- Automated LLM evaluation or prompt regression pipelines
- Worked on agent-based or multi-step AI workflows
- Debugged incorrect or hallucinated model outputs using structured methods
- Established quality metrics where ground truth was unclear or evolving
Zensar Technologies Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Zensar Technologies and has not been reviewed or approved by Zensar Technologies.
-
Retirement Support — A 401(k) with company match and immediate vesting is highlighted, with broad fund options available.
-
Wellbeing & Lifestyle Benefits — Work-life balance, remote/hybrid flexibility, and supportive teams are emphasized as strengths that enhance the overall experience.
-
Parental & Family Support — Inclusive policies such as parental leave and access to EAP and dependent medical coverage are promoted across regions.
Zensar Technologies Insights
What We Do
Zensar is a leading experience, engineering, and technology solutions company. We conceptualize, build, and manage digital products for Forbes Global 2000 clients across the hi-tech engineering, banking and financial services, insurance, manufacturing and consumer services verticals. With proven excellence across five core areas, including experience services, advanced engineering services, data engineering and analytics, foundation services, and application services, our solutions leverage industry-leading platforms to help our clients be competitive, agile, and disruptive while moving with velocity through change and opportunity. Zensar’s expansive ecosystem of 60+ technology partners, including Oracle, Salesforce, SAP, Guidewire, Automation Anywhere, Adobe, and UiPath, enables us to deliver comprehensive solutions to clients, facilitating seamless integration and allowing them to leverage cutting-edge technologies and tools for enhanced business outcomes. Zensar is part of the USD 4.4 billion RPG Group. With headquarters in Pune, India, our 10,500+ employees, representing over 50 nationalities, work from 30+ locations across North America, UK/Europe, and South Africa. Visit us at www.zensar.com








