AI Quality Engineer
Location:- Hyderabad
Work Mode:- WFO (5 days a week)
Role Type:- Contractual (3 months and extension would depend on project requirement and performance)
Key Responsibilities
- Test and validate AI-generated insights, recommendations, and decision-making workflows.
- Evaluate LLM and RAG systems for accuracy, relevance, consistency, factuality, and hallucinations.
- Validate retrieval quality, context relevance, grounding, and response quality in RAG systems.
- Test AI agents and autonomous workflows across functional, negative, and edge-case scenarios.
- Define AI evaluation criteria, test datasets, quality metrics, and validation processes.
- Perform regression testing for models, prompts, RAG configurations, and AI workflows.
- Collaborate with AI/ML engineers to identify issues and improve AI system quality.
Requirements
Required Skills
- Strong understanding of AI/ML and Generative AI testing.
- Hands-on experience testing LLM and RAG-based applications.
- Knowledge of LLM evaluation, hallucination detection, relevance, and response quality.
- Understanding of AI agents and recommendation systems.
- Strong analytical and problem-solving skills.
Good to Have
- Experience with PyTest and automated testing frameworks.
- Experience building automated AI evaluation and regression frameworks.
- Familiarity with tools such as RAGAS, DeepEval, LangSmith, or equivalent.
- Experience with CI/CD-based test automation, performance testing, or AI guardrails.
- Familiarity with cloud platforms (AWS/Azure/GCP) and observability tools.
Success Metrics
- High accuracy, relevance, and reliability of AI outputs.
- Strong evaluation coverage across critical AI workflows.
- Early detection and reduction of hallucinations and AI regressions.
- Reduced production AI quality issues.
- Increased confidence and trust in AI-generated insights and recommendations.
Skills Required
- Strong understanding of AI/ML and Generative AI testing
- Hands-on experience testing LLM and RAG-based applications
- Knowledge of LLM evaluation, hallucination detection, relevance, and response quality
- Understanding of AI agents and recommendation systems
- Strong analytical and problem-solving skills
- Experience with PyTest and automated testing frameworks
- Experience building automated AI evaluation and regression frameworks
- Familiarity with RAGAS, DeepEval, LangSmith, or equivalent tools
- Experience with CI/CD-based test automation, performance testing, or AI guardrails
- Familiarity with AWS, Azure, or GCP and observability tools
What We Do
Side transforms high-performing agents, teams, and independent brokerages into successful businesses and boutique brands that are 100% agent-owned. Headquartered in San Francisco, Side exclusively partners with the best agents, empowering them with proprietary technology and a premier support team so they can be more productive, grow their business, and focus on serving their clients.
Why Work With Us
Led by a team of experienced industry professionals and technology innovators, the Side team is constantly developing technology that improves agent productivity, legal compliance, marketing programs, and customer experience, to support the best-in-breed agents that we partner with.
Gallery






