Software Engineer in Test - AI

Posted Yesterday
Be an Early Applicant
Salt Lake City, UT, USA
In-Office
Mid level
HR Tech
The Role
Design, build, and maintain automated test frameworks and AI evaluation suites for LLMs, RAG, and AI-driven services. Validate responsible AI controls, define quality metrics, integrate tests into CI/CD, run functional/performance/reliability tests, investigate production issues using observability data, and promote testing standards across engineering and product teams.
Summary Generated by Built In

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.

Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.

Location: Salt Lake City, UT

In this role, you will build automation, evaluate AI-driven behavior, validate responsible AI controls, and identify risks early so our products are reliable, scalable, production-ready, and trusted by millions of users.

Key Responsibilities

  • Become a Subject Matter Expert on the AI platform and possess a deep understanding of system interactions, upstream/downstream dependencies, and data flows.
  • Design, develop, and maintain automated test frameworks for AI platform services, APIs, agents, prompt-based workflows, and RAG-enabled applications.
  • Develop automated AI evaluation suites and regression benchmarks that continuously measure model behavior and detect quality degradation before release.
  • Build and execute functional, integration, end-to-end, regression, performance, and reliability tests for AI-driven systems and services.
  • Define evaluation strategies, quality metrics, and acceptance criteria for AI-generated outputs, including accuracy, relevance, consistency, grounding, safety, and business value.
  • Validate responsible AI controls, permissions, guardrails, data handling practices, and business rules that protect customer trust.
  • Partner with Engineering, Product, and Support throughout the software lifecycle to drive risk-based testing strategies, influence quality to ensure production readiness.
  • Integrate automated testing, AI evaluations, and release validation into CI/CD pipelines.
  • Investigate, triage, and communicate defects, quality issues, and production incidents, driving root cause analysis and continuous improvement.
  • Utilize observability tools, logs, metrics, and traces to investigate production issues and perform root cause analysis.
  • Establish and promote testing standards, automation patterns, and AI quality practices that improve reliability and delivery speed.

Required Qualifications

  • 3+ years of experience in software testing, quality engineering, test automation, or software development.
  • Hands-on experience developing automated tests, testing tools, or quality frameworks using Python, Selenium and Playwright.
  • Experience testing APIs, microservices, distributed systems, backend services, or event-driven architectures.
  • Strong experience testing AI-enabled applications using technologies such as LLMs, LangChain, LangGraph, or similar platforms.
  • Hands-on experience evaluating AI-generated outputs using datasets, scoring rubrics, golden test sets, benchmarking frameworks, or automated quality checks.
  • Experience testing RAG systems, including retrieval quality, embeddings, grounding, citations, context accuracy, and streaming responses.
  • Experience working with SQL and NoSQL technologies, including relational, document, key-value, or vector databases.
  • Experience integrating automated testing into CI/CD pipelines and modern software delivery practices.
  • Experience investigating production issues and using incident and defect data to improve system quality and reliability.
  • Strong understanding of test automation, performance testing, risk-based testing, release validation, and responsible AI principles.
  • Excellent collaboration and communication skills, with the ability to clearly articulate quality risks and acceptance criteria.

Preferred Qualifications

  • Experience with AWS, Kubernetes, Docker, and cloud-native architectures.
  • Experience testing event-driven systems using Kafka or similar messaging platforms.
  • Familiarity with application security testing and OWASP Top 10 principles.
  • Proficiency in Python and at least one additional programming language such as Java, Ruby, or Go.
  • Experience with reliability engineering practices, including observability, production readiness reviews, incident analysis, SLOs, and continuous improvement.
  • Experience conducting performance, load, stress, scalability, or resilience testing.

Skills Required

  • 3+ years of experience in software testing, quality engineering, test automation, or software development
  • Hands-on experience developing automated tests, testing tools, or quality frameworks using Python
  • Hands-on experience with Selenium and Playwright
  • Experience testing APIs, microservices, distributed systems, backend services, or event-driven architectures
  • Strong experience testing AI-enabled applications using LLMs, LangChain, LangGraph, or similar platforms
  • Hands-on experience evaluating AI-generated outputs using datasets, scoring rubrics, golden test sets, benchmarking frameworks, or automated quality checks
  • Experience testing RAG systems including retrieval quality, embeddings, grounding, citations, and streaming responses
  • Experience working with SQL and NoSQL technologies, including relational, document, key-value, or vector databases
  • Experience integrating automated testing into CI/CD pipelines and modern software delivery practices
  • Experience investigating production issues and using incident and defect data to improve system quality and reliability
  • Strong understanding of test automation, performance testing, risk-based testing, release validation, and responsible AI principles
  • Excellent collaboration and communication skills, with ability to articulate quality risks and acceptance criteria
  • Experience with AWS, Kubernetes, Docker, and cloud-native architectures
  • Experience testing event-driven systems using Kafka or similar messaging platforms
  • Familiarity with application security testing and OWASP Top 10 principles
  • Proficiency in Python and at least one additional programming language such as Java, Ruby, or Go
  • Experience with reliability engineering practices, observability, production readiness reviews, incident analysis, SLOs, and continuous improvement
  • Experience conducting performance, load, stress, scalability, or resilience testing

O.C. Tanner Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about O.C. Tanner and has not been reviewed or approved by O.C. Tanner.

  • Retirement Support Retirement contributions are positioned as market‑leading with strong employer matching, and the retirement program is frequently highlighted as a standout part of the total package.
  • Strong & Reliable Incentives Bonuses and profit‑sharing are established components of total compensation, with periodic and consistent payouts contributing meaningful value beyond base pay.
  • Healthcare Strength Multiple medical plan options and employer‑supported health resources, including onsite services, indicate robust healthcare coverage complemented by wellness incentives and advisory support.

O.C. Tanner Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Salt Lake City, UT
1,300 Employees
Year Founded: 1927

What We Do

O.C. Tanner develops employee recognition strategies and rewards programs that help companies appreciate people who do great work. O.C. Tanner helps organizations inspire and appreciate great work. Thousands of clients globally use our cloud-based technology, tools, and awards to provide meaningful recognition for their employees. Learn more at www.octanner.com. ABOUT OUR PRODUCTS: Yearbook™ started a service award revolution. As the biggest innovation in service awards in 50 years, Yearbook has earned the right to be called a game-changer. Hundreds of thousands of recipients have loved the way Yearbook transforms service awards into unforgettable celebrations among friends at work. Check it out: http://www.octanner.com/products/celebrate-careers. Our popular Numeral™ awards capture career stages in trophies people love. Available in clear acrylic or metallic versions, Numerals can be customized to complement your brand or companion Yearbook. Explore awards: http://www.octanner.com/why-choose-us/awards-strategy When a person or team achieves outstanding results, big or small, it’s time to shine a spotlight on what they did and to reward their great work with an experience equal to their accomplishment. Our world-class performance recognition awards, programs, and fulfillment make it happen. Discover more about our performance and social platforms: http://www.octanner.com/products/performance-recognition Connect with us on... Twitter @octanner Facebook www.facebook.com/octannercompany Slideshare http://www.slideshare.net/octannercompany Instagram http://instagram.com/octannercompany YouTube http://www.youtube.com/user/octannercompany

Similar Jobs

In-Office
2 Locations
500 Employees
99K-124K Annually

Mondelēz International Logo Mondelēz International

Product Owner

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
United States
90000 Employees
140K-193K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
45K-85K Annually

HiBob Logo HiBob

People and Culture Partner

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
United States
1350 Employees
120K-150K Annually

Similar Companies Hiring

RethinkFirst Thumbnail
Telehealth • Software • Professional Services • Information Technology • HR Tech • Healthtech • Edtech
New York, NY
300 Employees
Empathy Thumbnail
Fintech • Healthtech • HR Tech • Information Technology • Financial Services • Telehealth
IL
200 Employees
Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account