Senior AI Quality Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Ontario, ON, CAN
Remote
100K-130K Annually
Senior level
Software
The Role
Own quality evaluation for LLM-driven and agentic products by building regression suites, golden datasets, scoring rubrics, and test harnesses. Detect hallucinations, tool-calling failures, prompt regressions, and unsafe behavior. Define reliability metrics, integrate evaluations into CI/CD, improve observability and tracing, contribute to API and end-to-end automation, and coach QA engineers on AI testing practices.
Summary Generated by Built In
Novara provides safety, operational risk management, and sustainability software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations manage operational risk proactively with a system they can configure to fit their operations, data that unifies risk and safety intelligence, and tools that increase workforce participation. Novara’s combination of software, training, and AI-powered tools put people and safety first while protecting critical operations.

Position Description:

AI is core to where our product is heading, and the quality bar for it has to be as high as anything else we ship. This role owns that bar. You will define what "working" means for non-deterministic systems, build the tooling to prove it, and give our teams the confidence to move faster on AI-powered features and agent-based products.

You will partner closely with product, engineering, and the broader QA team to bring rigor to how we test LLM-driven and agentic workflows, while also contributing to traditional automation coverage where it matters.

This is a hands-on IC role reporting into the QA Manager, with high visibility into AI initiatives across the company.

Responsibilities:

  • Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
  • Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
  • Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
  • Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
  • Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
  • Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
  • Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams

Knowledge, Experience, Requirements:

  • 6+ years in QA, SDET, or test automation, with real production automation shipping
  • Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
  • Prior experience in a shift-left, embedded QA model
  • Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
  • Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
  • API-first testing mindset, including REST and Postman or equivalent
  • Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
  • Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
  • Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
  • Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise

Preferred Qualifications:

    Not required, but they will stand out:

  • Experience with eval tooling such as Langfuse, Braintrust, LangSmith, Ragas, or DeepEval
  • Familiarity with RAG systems, vector stores, or tool-calling frameworks
  • Background in test data strategy for AI, including synthetic data generation
  • Exposure to Datadog or a similar observability platform

Tech Stack:

    TypeScript, Playwright, vitest, Postman, GitHub Actions, Jira/Xray, Datadog, AWS, plus emerging AI evaluation tooling.

Compensation:

    Annual Base Salary Range of 100k-130k CAD
    Annual Bonus Opportunity of 10%

As a growing company, Novara values its employees by supporting them with a full benefits package including Medical, Dental, Vision, Flexible Spending Accounts, PTO, Paid and Floating Holidays, 401k with Company match and immediate vesting, Company-funded Life Insurance, Employee Assistance Programs, and No-cost Mental Health Benefits.
 
About Novara
 
Novara provides safety and operational risk management software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations address operational risk proactively by unifying data, increasing workforce engagement, and proactively managing risk. Novara’s combination of training, software, and tools puts people and safety first while protecting critical operations.
 
Novara, a Providence Equity portfolio company, provides safety and operational risk management software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations address operational risk proactively by unifying data, increasing workforce engagement, and proactively managing risk. Novara’s combination of training, software, and tools puts people and safety first while protecting critical operations.
 
Novara launched January 1 2026, as an independent company, a spin-off of the Flex and RMC software businesses formerly part of KPA.
 
 
Don’t meet every job requirement? At Novara, we are dedicated to building a diverse, inclusive, and authentic workplace. Studies have shown that women and people of color are less likely to apply unless they meet every requirement. If you’re excited about the role but your past experience doesn’t align perfectly with every qualification, we still encourage you to apply! You might just be the right candidate for this or other roles.
 
Please note that we may use AI tools to assist in the initial screening of resumes to help identify qualified candidates more efficiently. All decisions are reviewed by a human recruiter, and no hiring determination is made solely by automated means. 
 
Novara is committed to providing equal opportunity in all of our employment practices, including selection, hiring, promotion, transfer, and compensation, to all qualified applicants and employees without regard to race, religion, religious dress/grooming, color, ethnicity, sex (including sex stereotyping), sexual orientation, gender identity or gender expression, national origin, ancestry, citizenship status, creed, uniform service member status, military or veteran status, marital status, pregnancy, breast-feeding and/or pregnancy-related conditions, age, protected medical condition, leave status, physical or mental disability, genetic characteristics, or any other legally-protected status in accordance with the requirements of all federal, state and local laws. In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification document form upon hire.
 
If you need assistance or an accommodation due to a disability, you may contact us at [email protected].
 
Please see our Candidate Privacy Notice Included Here 

Skills Required

  • 6+ years of experience in QA, SDET, or test automation with production automation experience
  • Hands-on experience testing LLM-based or agentic systems, including evaluations, LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
  • Experience working in a shift-left, embedded QA model
  • Experience with a modern automation stack such as Playwright or Cypress
  • Proficiency in a typed language, preferably TypeScript; Python is acceptable
  • Deep experience with Vitest, Jest, Pytest, or comparable test frameworks
  • Experience building custom test harnesses
  • API-first testing experience with REST and Postman or equivalent
  • Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
  • Working knowledge of CI/CD pipelines and integrating evaluations into them
  • Familiarity with cloud secret managers and handling sensitive test data in restore-from-production environments
  • Ability to reason about probabilistic systems, including variance, sample sizes, confidence, and flaky results
  • Experience with AI evaluation tools such as Langfuse, Braintrust, LangSmith, Ragas, or DeepEval
  • Familiarity with RAG systems, vector stores, or tool-calling frameworks
  • Experience with AI test data strategy or synthetic data generation
  • Exposure to Datadog or a similar observability platform
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
160 Employees

What We Do

Novara provides safety and operational risk management software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations address operational risk proactively by unifying data, increasing workforce engagement, and proactively managing risk. Novara’s combination of training, software, and tools puts people and safety first while protecting critical operations.

Similar Jobs

Atlassian Logo Atlassian

Software Engineer

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
Canada
11000 Employees
118K-154K Annually

Samsara Logo Samsara

Customer Success Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Canada
4000 Employees
78K-101K Annually

Sprout Social Logo Sprout Social

Product Manager

Marketing Tech • Social Media • Software • Analytics • Business Intelligence
Easy Apply
Remote or Hybrid
Canada
1400 Employees
145K-218K Annually

Affirm Logo Affirm

Staff Software Engineer

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
Canada
2200 Employees
181K-241K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account