Senior Site Reliability Engineer – Unified Observability

Posted 2 Days Ago
Be an Early Applicant
Atlanta, GA, USA
In-Office
Senior level
Fintech • Information Technology • Payments • Software
The Role
Lead design and implementation of a unified enterprise observability platform across Azure, GCP, Kubernetes and hybrid environments. Define observability standards (monitoring, logging, tracing, SLIs/SLOs), build dashboards, integrate with CI/CD and ServiceNow, drive incident prevention/response automation, and promote AI-driven analytics. Influence architecture, mentor engineers, and partner cross-functionally to improve platform reliability and operational readiness.
Summary Generated by Built In

About NCR VOYIX

NCR Voyix Corporation (NYSE: VYX) is a global platform-powered leader in unified commerce for shopping and dining. Combining a flexible, intelligent platform with end-to-end payments capabilities and services developed through its deep industry experience, NCR Voyix empowers retailers and restaurants to accelerate new possibilities for their operations, experiences and business outcomes. NCR Voyix is headquartered in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.
Position Overview

We are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer Unified Observability initiative. This strategic role will be responsible for building and evolving a unified enterprise observability platform that delivers end-to-end visibility across NCR Voyix Restaurants, Retail, and Payments environments.

The ideal candidate will bring 10+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, or related disciplines, with a proven track record of driving enterprise-scale observability, reliability, and operational excellence. This individual must be comfortable operating across organizational boundaries and partnering closely with Product Engineering, Infrastructure, Security, Operations, Architecture, and Executive Leadership teams to establish a comprehensive observability strategy and improve platform resilience.

This role will serve as a key technical leader responsible for defining standards, influencing architecture decisions, and enabling proactive operations through unified monitoring, telemetry, automation, and AI-driven insights.

Key Responsibilities
  • Lead the architecture, design, implementation, and continuous improvement of enterprise observability solutions across Azure, Google Cloud Platform (GCP), Kubernetes, and hybrid environments.
  • Establish and drive enterprise observability standards for monitoring, logging, distributed tracing, telemetry, and operational analytics.
  • Develop and maintain executive, operational, and engineering dashboards that provide real-time visibility into infrastructure, applications, platform health, customer experience, and business transactions.
  • Define, evangelize, and implement reliability frameworks including SLIs, SLOs, error budgets, operational KPIs, and service health metrics.
  • Partner cross-functionally with Engineering, Infrastructure, Security, Product, and Operations teams to identify reliability risks and drive operational excellence initiatives.
  • Lead efforts to improve incident prevention, detection, response, and recovery through intelligent alerting, automation, event correlation, and observability best practices.
  • Integrate observability capabilities with ServiceNow, CI/CD pipelines, automation frameworks, and enterprise operational workflows.
  • Influence technical strategy and roadmap decisions related to reliability engineering, platform observability, and operational readiness.
  • Support and drive enterprise initiatives involving AI-driven observability, predictive analytics, anomaly detection, and event intelligence.
  • Mentor engineers and serve as a subject matter expert for observability, reliability engineering, and cloud-native operations.
  • Establish governance, adoption, and best practices across multiple product and engineering teams to ensure consistent observability standards enterprise-wide.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 10+ years of experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, or related technical disciplines.
  • Demonstrated success designing and operating observability platforms in large-scale enterprise environments.
  • Deep expertise with Kubernetes platforms, including AKS and GKE.
  • Strong experience with Azure and Google Cloud Platform services and architectures.
  • Hands-on experience with enterprise observability tools such as Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, or similar platforms.
  • Advanced knowledge of monitoring, logging, telemetry collection, distributed tracing, and observability engineering principles.
  • Experience defining and operationalizing SLIs, SLOs, error budgets, reliability metrics, and service health frameworks.
  • Strong automation and Infrastructure as Code expertise using Terraform and related tools.
  • Proficiency developing automation solutions using Python, Go, PowerShell, or similar languages.
  • Experience integrating observability solutions into CI/CD pipelines and modern DevOps workflows.
  • Proven ability to influence technical direction and collaborate effectively with stakeholders across Engineering, Product, Infrastructure, Security, and Operations organizations.
  • Strong communication, leadership, and stakeholder management skills with the ability to translate technical concepts for both technical and business audiences.
Preferred Qualifications
  • Experience leading enterprise observability transformations or platform modernization initiatives.
  • Experience with AI Ops, event correlation, operational analytics, and predictive monitoring capabilities.
  • Knowledge of ServiceNow integrations and ITSM/ITOM processes.
  • Experience supporting highly available, customer-facing SaaS platforms at scale.
  • One or more cloud certifications (Azure, Google Cloud, Kubernetes, or related technologies).
  • Previous experience serving as a technical lead, mentor, or architect within a reliability engineering organization.

Offers of employment are conditional upon passage of screening criteria applicable to the job

EEO Statement

Integrated into our shared values is NCR Voyix’s commitment to equal employment opportunity.  All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law.  NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential.  We believe in understanding and respecting differences among all people.  Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.

Statement to Third Party Agencies
To ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes

“When applying for a job, please make sure to only open emails that you will receive during your application process that come from a @ncrvoyix.com email domain.”

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience
  • 10+ years experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, or related disciplines
  • Designing and operating enterprise-scale observability platforms
  • Deep expertise with Kubernetes platforms including AKS and GKE
  • Strong experience with Azure and Google Cloud Platform services and architectures
  • Hands-on experience with observability tools such as Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic or similar
  • Advanced knowledge of monitoring, logging, telemetry collection, distributed tracing, and observability engineering principles
  • Experience defining and operationalizing SLIs, SLOs, error budgets, reliability metrics, and service health frameworks
  • Strong automation and Infrastructure as Code expertise using Terraform
  • Proficiency developing automation solutions using Python, Go, PowerShell, or similar languages
  • Experience integrating observability solutions into CI/CD pipelines and modern DevOps workflows
  • Ability to influence technical direction and collaborate with Engineering, Product, Infrastructure, Security, and Operations stakeholders
  • Strong communication, leadership, and stakeholder management skills
  • Experience leading enterprise observability transformations or platform modernization initiatives
  • Experience with AI Ops, event correlation, operational analytics, and predictive monitoring
  • Knowledge of ServiceNow integrations and ITSM/ITOM processes
  • Experience supporting highly available, customer-facing SaaS platforms at scale
  • One or more cloud certifications (Azure, Google Cloud, Kubernetes, or related technologies)
  • Previous experience serving as a technical lead, mentor, or architect within a reliability engineering organization

NCR Corporation Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NCR Corporation and has not been reviewed or approved by NCR Corporation.

  • Healthcare Strength Healthcare coverage is described as comprehensive, including medical, dental, and vision, alongside HSA/HRA funding and an employee assistance program. This breadth increases the perceived baseline value of the total rewards package even when pay satisfaction varies.
  • Retirement Support Retirement support includes a 401(k) match structure and an employee stock purchase program with a stated discount. These elements provide longer-term wealth-building mechanisms beyond base salary.
  • Leave & Time Off Breadth Time-off provisions include paid vacation, holidays (including floating days), sick time, and defined maternity and paternity leave. This breadth can improve the overall rewards experience for those who can fully utilize the leave policies.

NCR Corporation Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Atlanta, GA
36,000 Employees
Year Founded: 1884

What We Do

Shaping the future for 135 years, NCR is the world’s enterprise technology leader for restaurants, retailers and banks. The #1 global POS software provider for retail and hospitality, and the #1 provider of multi-vendor ATM software, we create software, hardware and services that run the enterprise from back office to the front end and everything in between for our clients.

Similar Jobs

SambaSafety Logo SambaSafety

Business Intelligence Analyst

Insurance • Logistics • Software • Transportation • Business Intelligence
Remote or Hybrid
United States
300 Employees
100K-110K Annually

PwC Logo PwC

Cybersecurity - Privileged Access Management - Sr Associate

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
21 Locations
370000 Employees
77K-202K Annually

PwC Logo PwC

Chief Financial Officer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Remote or Hybrid
33 Locations
370000 Employees
123K-123K Annually

PwC Logo PwC

Healthcare Provider, Business Operations - Senior Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
39 Locations
370000 Employees
124K-280K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account