AI is core to where our product is heading, and the quality bar for it has to be as high as anything else we ship. This role owns that bar. You will define what "working" means for non-deterministic systems, build the tooling to prove it, and give our teams the confidence to move faster on AI-powered features and agent-based products.
You will partner closely with product, engineering, and the broader QA team to bring rigor to how we test LLM-driven and agentic workflows, while also contributing to traditional automation coverage where it matters.
This is a hands-on IC role reporting into the QA Manager, with high visibility into AI initiatives across the company.
Responsibilities:
- Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
- Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
- Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
- Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
- Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
- Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
- Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams
Knowledge, Experience, Requirements:
- 6+ years in QA, SDET, or test automation, with real production automation shipping
- Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
- Prior experience in a shift-left, embedded QA model
- Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
- Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
- API-first testing mindset, including REST and Postman or equivalent
- Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
- Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
- Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
- Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise
Preferred Qualifications:
- Experience with eval tooling such as Langfuse, Braintrust, LangSmith, Ragas, or DeepEval
- Familiarity with RAG systems, vector stores, or tool-calling frameworks
- Background in test data strategy for AI, including synthetic data generation
- Exposure to Datadog or a similar observability platform
Not required, but they will stand out:
Tech Stack:
TypeScript, Playwright, vitest, Postman, GitHub Actions, Jira/Xray, Datadog, AWS, plus emerging AI evaluation tooling.
Compensation:
Skills Required
- 6+ years of experience in QA, SDET, or test automation with production automation experience
- Hands-on experience testing LLM-based or agentic systems, including evaluations, LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
- Experience working in a shift-left, embedded QA model
- Experience with a modern automation stack such as Playwright or Cypress
- Proficiency in a typed language, preferably TypeScript; Python is acceptable
- Deep experience with Vitest, Jest, Pytest, or comparable test frameworks
- Experience building custom test harnesses
- API-first testing experience with REST and Postman or equivalent
- Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
- Working knowledge of CI/CD pipelines and integrating evaluations into them
- Familiarity with cloud secret managers and handling sensitive test data in restore-from-production environments
- Ability to reason about probabilistic systems, including variance, sample sizes, confidence, and flaky results
- Experience with AI evaluation tools such as Langfuse, Braintrust, LangSmith, Ragas, or DeepEval
- Familiarity with RAG systems, vector stores, or tool-calling frameworks
- Experience with AI test data strategy or synthetic data generation
- Exposure to Datadog or a similar observability platform
What We Do
Novara provides safety and operational risk management software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations address operational risk proactively by unifying data, increasing workforce engagement, and proactively managing risk. Novara’s combination of training, software, and tools puts people and safety first while protecting critical operations.









