O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.
Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.
Location: Salt Lake City, UT
As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.
Key Responsibilities:
- Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
- Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.
- Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
- Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
- Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
- Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.
- Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.
- Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
- Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
- Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
- Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.
Required Qualifications
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
- Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
- Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
- Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
- Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
- Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
- Strong knowledge of AWS and Kubernetes in production environments.
- Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.
- Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
- Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews.
Preferred Qualifications
- Experience leading distributed or globally dispersed engineering teams.
- Experience with multiple cloud providers or cloud-agnostic platform architectures.
- Familiarity with security, compliance, governance, and operational risk management frameworks.
- Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
- Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
- Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.
- Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.
Skills Required
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines
- 2+ years in a technical leadership or people management role
- Proven experience leading teams responsible for production operations, reliability engineering, and incident management
- Experience designing and implementing SRE practices and operational maturity initiatives
- Experience operating large-scale, customer-facing SaaS platforms with high availability and scalability requirements
- Hands-on experience with observability platforms such as OpenTelemetry, Datadog, or Coralogix
- Strong knowledge of AWS in production environments
- Strong knowledge of Kubernetes in production environments
- Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles
- Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support
- Experience developing engineering roadmaps, defining team objectives, and aligning reliability investments with business priorities
- Experience with production incident response, root cause analysis, and blameless post-incident reviews
- Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning
- Ability to collaborate with global engineering teams in a follow-the-sun support model
- Ownership of on-call programs, incident management practices, and operational health metrics
- Experience leading distributed or globally dispersed engineering teams
- Experience with multiple cloud providers or cloud-agnostic platform architectures
- Familiarity with security, compliance, governance, and operational risk management frameworks
- Proficiency with Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and performance monitoring tools
- Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora
- Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS
- Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics
O.C. Tanner Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about O.C. Tanner and has not been reviewed or approved by O.C. Tanner.
-
Retirement Support — Retirement contributions are positioned as market‑leading with strong employer matching, and the retirement program is frequently highlighted as a standout part of the total package.
-
Strong & Reliable Incentives — Bonuses and profit‑sharing are established components of total compensation, with periodic and consistent payouts contributing meaningful value beyond base pay.
-
Healthcare Strength — Multiple medical plan options and employer‑supported health resources, including onsite services, indicate robust healthcare coverage complemented by wellness incentives and advisory support.
O.C. Tanner Insights
What We Do
O.C. Tanner develops employee recognition strategies and rewards programs that help companies appreciate people who do great work. O.C. Tanner helps organizations inspire and appreciate great work. Thousands of clients globally use our cloud-based technology, tools, and awards to provide meaningful recognition for their employees. Learn more at www.octanner.com. ABOUT OUR PRODUCTS: Yearbook™ started a service award revolution. As the biggest innovation in service awards in 50 years, Yearbook has earned the right to be called a game-changer. Hundreds of thousands of recipients have loved the way Yearbook transforms service awards into unforgettable celebrations among friends at work. Check it out: http://www.octanner.com/products/celebrate-careers. Our popular Numeral™ awards capture career stages in trophies people love. Available in clear acrylic or metallic versions, Numerals can be customized to complement your brand or companion Yearbook. Explore awards: http://www.octanner.com/why-choose-us/awards-strategy When a person or team achieves outstanding results, big or small, it’s time to shine a spotlight on what they did and to reward their great work with an experience equal to their accomplishment. Our world-class performance recognition awards, programs, and fulfillment make it happen. Discover more about our performance and social platforms: http://www.octanner.com/products/performance-recognition Connect with us on... Twitter @octanner Facebook www.facebook.com/octannercompany Slideshare http://www.slideshare.net/octannercompany Instagram http://instagram.com/octannercompany YouTube http://www.youtube.com/user/octannercompany





.jpg)


