Senior Manager, Data & Storage Reliability Engineering

Posted Yesterday
Be an Early Applicant
Dublin, IRL
Hybrid
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
We're putting AI to work for people.
The Role
Leads Data and Storage Reliability Engineering for large-scale SaaS environments. Responsibilities include reliability strategy, observability, diagnostics, automation, capacity planning, incident root-cause analysis, resilience validation, risk reduction, and platform improvement. Manages and develops engineering teams, partners with production and software engineering groups, supports customer-critical escalations, and converts operational insights into durable engineering solutions.
Summary Generated by Built In
Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

 

Job Description

ServiceNow is seeking an experienced Sr Manager for our Data & Storage Reliability Engineering team. 

This leader will drive engineering excellence across prevention engineering, reliability engineering, observability, incident learnings, diagnostics, automation, capacity planning, and platform risk reduction. The role requires deep technical expertise in database and distributed systems architectures, large-scale SaaS production environments, and customer-facing operations, combined with strong people leadership and execution rigor. 

The ideal candidate has experience leading high-performing engineering teams responsible for identifying recurring production patterns, converting incident and escalation learnings into durable engineering improvements, strengthening observability, and building sustainable solutions that improve platform resilience at scale. 

You should have experience with large-scale web applications, database platforms, distributed systems, Linux-based production environments, and a strong problem-solving mindset for reliability, automation, diagnostics, observability, and prevention. Qualified candidates will be responsible for leading a team of highly skilled engineers that push the limits of scalability, resiliency, and operational excellence. 

Do you 

  • Have experience leading teams of engineers and developing people? 
  • Enjoy problem solving and using an analytical mindset to understand why systems fail and how to prevent repeat issues? 
  • Have a technical background in roles including database engineering, reliability engineering, systems/cloud engineering, SRE, DevOps, or production engineering? 
  • Know Linux operating systems, databases, observability, diagnostics, and production troubleshooting well enough to guide engineers through complex investigations? 
  • Have an attitude of continuous improvement and a passion for removing inefficient, repetitive, or reactive processes through automation and engineering prevention? 

Answer 'yes' to these questions and we want to hear from you. Hit the Apply button and let's have a chat about the role and your skills and experiences. 

Let’s start with the role 

As a Sr Manager of the Data & Storage Reliability Engineering team your responsibilities will be: 

  • Define and execute the team-level strategy for prevention engineering, reliability, observability, resilience, and operational risk reduction across large-scale production environments. 
  • Lead initiatives that turn production signals, incident learnings, customer escalations, migration outcomes, and platform telemetry into durable engineering improvements. 
  • Partner closely with SWAT and Customer & Production Engineering to establish a continuous feedback loop between production operations and platform improvement. 
  • Drive improvements in observability, diagnostics, automation, reliability reviews, resiliency validation, migration readiness, and engineering guardrails. 
  • Identify recurring failure patterns, reliability risks, observability gaps, operational inefficiencies, scalability constraints, and performance bottlenecks, and drive action to reduce future customer impact. 
  • Establish reliability, resilience, observability, automation, and prevention goals for critical database and storage services. 
  • Champion proactive monitoring, production analytics, and automation to improve operational health and reduce repetitive manual work. 
  • Lead deep root cause analysis and ensure sustainable corrective actions are implemented for recurring issues and customer-impacting events. 
  • Partner with engineering leaders to influence database, storage, reliability, observability, and platform architecture priorities based on production evidence. 
  • Build and develop a world-class team of reliability, prevention, observability, and platform engineers. 
  • Own team management, recruitment, career development, objective setting, project prioritization, onboarding, and performance reviews. 
  • Manage an engineering team that supports production-facing work, including on-call or escalation participation where required. 
  • Drive a culture of intolerance for repetitive manual activities by promoting automation, self-service diagnostics, guardrails, and scalable engineering practices. 
  • Drive initiatives with partner teams to improve the reliability, resilience, scalability, and operational efficiency of the ServiceNow application and platform. 
  • Act as part of the escalation and crisis management ecosystem by helping convert immediate recovery learnings into sustainable engineering prevention. 
  • Analyze and evaluate existing processes to drive continuous improvement, operational efficiency, and prevention-oriented engineering practices. 
  • Provide training, documentation, dashboards, playbooks, and support to partner teams that interface with the Data & Storage Reliability Engineering team. 
  • Onboard new hires, new technologies, new systems, and new automations into the team to enable successful execution and scale. 

Qualifications

  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving — using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function. 

  • 10+ years of experience in database engineering, reliability engineering, distributed systems, platform engineering, infrastructure engineering, production engineering, or large-scale SaaS platform operations. 

  • 4+ years of engineering leadership experience, including leading engineers and cross-functional or distributed teams. 

  • Experience leading Reliability Engineering, Database Engineering, Platform Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical teams. 

  • Strong expertise in database technologies, operating system performance, distributed systems, cloud-native architectures, and large-scale production environments. 

  • Solid understanding of reliability engineering, observability, diagnostics, root cause analysis, capacity planning, scalability engineering, resiliency, automation, and operational excellence. 

  • Experience translating production insights, customer escalations, incident learnings, platform telemetry, and recurring operational challenges into prioritized engineering work. 

  • Experience designing and improving observability, diagnostics, reliability reviews, migration readiness checks, resiliency validation, automation, and engineering guardrails. 

  • Strong understanding of tuning and troubleshooting across database, operating system, storage, network, and application layers. 

  • Proven experience identifying and resolving complex reliability, scalability, performance, efficiency, and operational bottlenecks in large-scale distributed environments. 

  • Experience leading customer-critical investigations involving reliability, capacity, performance, scalability, resilience, or operational risk challenges. 

  • Experience leveraging observability and telemetry platforms to analyze system behavior and drive platform improvements. 

  • Experience partnering with software engineering, infrastructure, production operations, and escalation organizations to improve platform reliability, scalability, efficiency, and production readiness. 

  • Experience driving engineering initiatives through data, metrics, incident learnings, telemetry, benchmarking, and measurable outcomes. 

  • Strong communication, stakeholder management, and leadership skills. 

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. 

Desired Skills 

  • Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads. 

  • Experience growing Reliability Engineering, Database Engineering, Platform Engineering, Performance Engineering, Scalability Engineering, or Production Engineering teams. 

  • Experience with observability platforms, telemetry systems, diagnostics frameworks, production analytics, reliability scorecards, engineering metrics, and impact reporting. 

  • Experience with reliability reviews, resiliency validation, migration readiness, workload simulation, capacity forecasting, prevention programs, and operational risk reduction frameworks. 

  • Experience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, operational efficiency, and engineering productivity. 

  • Understanding of distributed systems architecture, cloud platform operations, Linux-based production environments, and hyperscale environments. 

  • Experience contributing to platform architecture, database strategy, reliability investments, scalability roadmaps, and long-term engineering improvements. 

  • Experience with performance testing, benchmarking, workload simulation, and capacity modeling as part of broader reliability and prevention engineering programs. 

  • Experience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms. 

  • Familiarity with ServiceNow platform architecture and large-scale SaaS operations. 

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity,  veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.  

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Skills Required

  • Experience integrating or critically evaluating AI in work processes, decision-making, or problem-solving
  • 10+ years of experience in database engineering, reliability engineering, distributed systems, platform engineering, infrastructure engineering, production engineering, or large-scale SaaS operations
  • 4+ years of engineering leadership experience leading engineers and cross-functional or distributed teams
  • Experience leading reliability, database, platform, infrastructure, production, performance, or related technical teams
  • Expertise in database technologies, operating system performance, distributed systems, cloud-native architectures, and large-scale production environments
  • Understanding of reliability engineering, observability, diagnostics, root-cause analysis, capacity planning, scalability engineering, resiliency, automation, and operational excellence
  • Experience translating production insights, customer escalations, incident learnings, telemetry, and operational challenges into prioritized engineering work
  • Experience designing and improving observability, diagnostics, reliability reviews, migration readiness, resiliency validation, automation, and engineering guardrails
  • Understanding of tuning and troubleshooting database, operating system, storage, network, and application layers
  • Experience resolving complex reliability, scalability, performance, efficiency, and operational bottlenecks in distributed environments
  • Experience leading customer-critical investigations involving reliability, capacity, performance, scalability, resilience, or operational risk
  • Experience using observability and telemetry platforms to analyze system behavior and drive platform improvements
  • Experience partnering with software engineering, infrastructure, production operations, and escalation organizations
  • Experience driving engineering initiatives through data, metrics, incident learnings, telemetry, benchmarking, and measurable outcomes
  • Strong communication, stakeholder management, and leadership skills
  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
  • Experience operating enterprise database and storage platforms supporting mission-critical workloads
  • Experience growing reliability, database, platform, performance, scalability, or production engineering teams
  • Experience with observability platforms, telemetry systems, diagnostics frameworks, production analytics, reliability scorecards, engineering metrics, and impact reporting
  • Experience with reliability reviews, resiliency validation, migration readiness, workload simulation, capacity forecasting, prevention programs, and operational risk reduction
  • Experience using AI for anomaly detection, forecasting, incident analysis, prioritization, operational efficiency, and engineering productivity
  • Understanding of distributed systems architecture, cloud platform operations, Linux production environments, and hyperscale environments
  • Experience contributing to platform architecture, database strategy, reliability investments, scalability roadmaps, and long-term engineering improvements
  • Experience with performance testing, benchmarking, workload simulation, and capacity modeling
  • Experience with MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms
  • Familiarity with ServiceNow platform architecture and large-scale SaaS operations

What the Team is Saying

Shanequa
Katya
Suzanne
Alexander
Jaime
Pat
Brady
Hasan
Jamil
Viviana

ServiceNow Compensation & Benefits Highlights

  • Healthcare Strength Health coverage is described as comprehensive with multiple plan choices and strong perceived coverage, alongside mental-health resources and wellbeing support. Company materials and employer-verified summaries also note inclusive care options and supportive programs.
  • Parental & Family Support Parental leave and family-planning support are characterized as generous, with fully paid leave and resources such as fertility, caregiving, and adoption assistance. Backup care and other family-focused programs are also highlighted as part of the package.
  • Equity Value & Accessibility Equity components like RSUs and an employee stock purchase plan are presented as meaningful, widely available parts of total rewards. Many role and benefits overviews emphasize equity’s role in boosting overall compensation alongside bonuses.

ServiceNow Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
29,000 Employees
Year Founded: 2004

What We Do

As the AI platform for business transformation, we're putting AI to work across organizations — freeing people for work that matters. Making old tech work with new tech. Reaching across departments, from the front office to the back office and every office in between. Our ambition? To become the AI defining enterprise software company of the 21st century (or "AI DESCO21C," as we like to call it). With more than 8,400+ customers, we serve approximately 90% of the Fortune 500®, and we're proud to be a Fortune 100 Best Companies to Work For® and World's Most Admired Companies™. Explore your future career with us, visit www.careers.servicenow.com From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Why Work With Us

By joining ServiceNow, you are part of an ambitious team of change-makers who have a restless curiosity and a drive for ingenuity. We're committed to helping our people do their best work and live their best lives so we can fulfill our purpose together. At the fastest-growing enterprise software company, you can grow your career faster.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

ServiceNow Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

At ServiceNow, we lead with flexibility and trust. For some, home is the primary workplace. For those who come into a ServiceNow workplace, you are empowered to make team-guided and individual-led decisions on how and when you use the workplace.

Typical time on-site: Flexible
Company Office Image
HQSanta Clara, CA
Heredia
Ciudad de México
District of Columbia
Osaka
Aarhus, DK
Aarhus, DK
Company Office Image
Addison, TX
Amsterdam, North Holland
Atlanta, Georgia
Auckland, Auckland
Bangkok, Bangkok
Bengaluru, Karnataka
Bengaluru, Karnataka
Berlin, Berlin
Brasília, Federal District
Brisbane, Queensland
Brussels, BE
Cairo, Cairo Governorate
Canberra, Australian Capital Territory
Charlottesville, Virginia
Company Office Image
Chicago, IL
Deerfield, Illinois
Company Office Image
Denver, CO
Dubai, Dubai
Dublin, Leinster
Düsseldorf, Nordrhein-Westfalen
Frankfurt am Main, Hesse
Goteborg, Västra Götaland County
Gurugram, Haryana
Hamburg, Hamburg
Hanyang, Seoul
Helsinki, Uusimaa
Hong Kong, Hong Kong
Houston, TX
Hyderabad, Telangana
Issy-les-Moulineaux, Île-de-France
Johannesburg, Gauteng
Kirkland, WA
Lausanne, Vaud
Lille, Hauts de France
London, England
London, England
Madrid, Community of Madrid
Melbourne, Victoria
Milano, Lombardia
Milwaukee, WI
Company Office Image
Montréal, QC
Mumbai, Maharashtra
Munich, Bavaria
Company Office Image
New York, NY
Opfikon, Zürich
Orlando, FL
Oslo, Oslo
Perth, Western Australia
Petah Tikva, Central District
Company Office Image
Pleasanton, CA
Riyadh, Riyadh Province
Rome, Lazio
Company Office Image
San Diego, CA
San Francisco, Heredia
Company Office Image
San Francisco, CA
São Paulo, SP
Singapore, SG
Solna, Stockholm County
Sydney, New South Wales
Tokyo, Tokyo
Toronto, Ontario
Vancouver, British Columbia
Company Office Image
Vienna, VA
Vienna, AT
Company Office Image
Waltham, MA
Washington, DC
Wellington, Wellington
West Palm Beach, Florida
Learn more

Similar Jobs

ServiceNow Logo ServiceNow

Senior Director, CEG EMEA Field Ops Lead

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Dublin, IRL
29000 Employees

ServiceNow Logo ServiceNow

Development Manager

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Dublin, IRL
29000 Employees

ServiceNow Logo ServiceNow

Senior Employee Relations (ER) Partner

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Dublin, IRL
29000 Employees

ServiceNow Logo ServiceNow

Staff Software Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Dublin, IRL
29000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account