Sr. Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Salt Lake City, UT, USA
In-Office
Senior level
HR Tech
The Role
Lead reliability engineering for a large-scale SaaS platform by building automation, observability, and self-healing systems. Drive incident response, triage, SLOs, and blameless postmortems. Partner with engineering, product, and support to embed reliability, run follow-the-sun operations, participate in on-call rotations, and reduce operational toil through tooling and automation.
Summary Generated by Built In

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.

Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.

Location: Salt Lake City, UT

As a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage software engineering, automation, and cloud-native technologies to build and operate highly available, scalable systems that serve millions of users. We're looking for someone who is passionate about reliability engineering, continuous improvement, and building self-healing platforms that enable development teams to move faster while delivering exceptional customer experiences.

Key Responsibilities

  • Improve the availability, scalability, and performance of cloud-native applications through automation, monitoring, and engineering best practices.
  • Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools. Establish standards for metrics, logs, traces, and service-level objectives (SLOs) that enable proactive issue detection and resolution.
  • Lead production triage efforts, rapidly diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews to improve system resilience and reduce recurring issues.
  • Partner with Engineering, Support and Product teams to embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Champion a reliability-first engineering culture by establishing automation standards, monitoring best practices, shift-left quality approaches and shared ownership models that proactive improve resilience, reduce operational risk, and protect the availability of business-critical services.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective handoffs, and shared ownership of production services.
  • Participate in an on-call rotation focused on maintaining service health, reducing operational toil, improving alert quality, and automating repetitive operational tasks.

 Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles, with a strong background in production triage, incident response, and operational excellence.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements.
  • Proficiency in Go, Python, Java, or similar programming languages, with demonstrated experience building automation, production tooling, and reliability-focused engineering solutions.
  • Deep experience with modern Infrastructure-as-Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes.
  • Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms.
  • Strong knowledge of AWS services and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, and distributed tracing for complex systems.
  • Ability to partner effectively with software engineering and testing teams to design reliable systems, improve application performance, and strengthen quality practices across the software development lifecycle.
  • Comfortable with participating in on-call rotations and handling high-pressure environments.

Bonus Qualifications:

  • Experience with multiple cloud or cloud-agnostic environments.
  • Familiarity with security, compliance, and governance frameworks
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.

Skills Required

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles
  • Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements
  • Proficiency in Go, Python, Java, or similar programming languages
  • Experience with Infrastructure-as-Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes
  • Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms
  • Strong knowledge of AWS services and Kubernetes in production environments
  • Deep understanding of monitoring, logging, and distributed tracing for complex systems
  • Ability to partner effectively with software engineering and testing teams to design reliable systems and improve application performance
  • Comfortable participating in on-call rotations and handling high-pressure environments
  • Experience with multiple cloud or cloud-agnostic environments
  • Familiarity with security, compliance, and governance frameworks
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS

O.C. Tanner Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about O.C. Tanner and has not been reviewed or approved by O.C. Tanner.

  • Retirement Support Retirement contributions are positioned as market‑leading with strong employer matching, and the retirement program is frequently highlighted as a standout part of the total package.
  • Strong & Reliable Incentives Bonuses and profit‑sharing are established components of total compensation, with periodic and consistent payouts contributing meaningful value beyond base pay.
  • Healthcare Strength Multiple medical plan options and employer‑supported health resources, including onsite services, indicate robust healthcare coverage complemented by wellness incentives and advisory support.

O.C. Tanner Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Salt Lake City, UT
1,300 Employees
Year Founded: 1927

What We Do

O.C. Tanner develops employee recognition strategies and rewards programs that help companies appreciate people who do great work. O.C. Tanner helps organizations inspire and appreciate great work. Thousands of clients globally use our cloud-based technology, tools, and awards to provide meaningful recognition for their employees. Learn more at www.octanner.com. ABOUT OUR PRODUCTS: Yearbook™ started a service award revolution. As the biggest innovation in service awards in 50 years, Yearbook has earned the right to be called a game-changer. Hundreds of thousands of recipients have loved the way Yearbook transforms service awards into unforgettable celebrations among friends at work. Check it out: http://www.octanner.com/products/celebrate-careers. Our popular Numeral™ awards capture career stages in trophies people love. Available in clear acrylic or metallic versions, Numerals can be customized to complement your brand or companion Yearbook. Explore awards: http://www.octanner.com/why-choose-us/awards-strategy When a person or team achieves outstanding results, big or small, it’s time to shine a spotlight on what they did and to reward their great work with an experience equal to their accomplishment. Our world-class performance recognition awards, programs, and fulfillment make it happen. Discover more about our performance and social platforms: http://www.octanner.com/products/performance-recognition Connect with us on... Twitter @octanner Facebook www.facebook.com/octannercompany Slideshare http://www.slideshare.net/octannercompany Instagram http://instagram.com/octannercompany YouTube http://www.youtube.com/user/octannercompany

Similar Jobs

Remote or Hybrid
United States
1750 Employees

ServiceTitan Logo ServiceTitan

Senior Site Reliability Engineer

Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Remote or Hybrid
US
2760 Employees
138K-221K Annually

Microsoft Logo Microsoft

Senior Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office or Remote
2 Locations
206870 Employees
120K-261K Annually

Akamai Technologies Logo Akamai Technologies

Senior Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
121K-219K Annually

Similar Companies Hiring

RethinkFirst Thumbnail
Telehealth • Software • Professional Services • Information Technology • HR Tech • Healthtech • Edtech
New York, NY
300 Employees
Empathy Thumbnail
Fintech • Healthtech • HR Tech • Information Technology • Financial Services • Telehealth
IL
200 Employees
Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account