Manager, Network Reliability and Resiliency

Posted 2 Hours Ago
Be an Early Applicant
Toronto, ON, CAN
Hybrid
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
We're putting AI to work for people.
The Role
Leads a team of network reliability engineers responsible for production network services supporting a cloud platform. Manages hiring, coaching, priorities, on-call practices, incident response, customer escalations, and operational readiness. Applies SRE practices, observability, automation, error budgets, and post-incident improvement to increase availability and reduce operational toil. Partners with engineering, SRE, security, infrastructure, and support teams while guiding complex troubleshooting and risk-based operational decisions.
Summary Generated by Built In
Description de l'entreprise

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

 

Description du poste

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status.  Employment is contingent upon successful completion and maintenance of the required screening.

What you get to do in this role:

We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow's cloud platform. This is a technical people-manager role. You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement.

You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent. While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions.

Lead and develop the team

  • Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development.
  • Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments.
  • Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events.
  • Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed.

Provide technical and incident leadership

  • Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts.
  • Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service.
  • Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions.
  • Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact.

Improve reliability through SRE practices

  • Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services.
  • Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk.
  • Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently.
  • Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover.

Embed automation in daily operations

  • Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated.
  • Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency.
  • Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models.
  • Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome.

Partner across the organization

  • Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability.
  • Represent the team's technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders.
  • Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership.
  • Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback.

Qualifications

To be successful in this role you have:

  • Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations.
  • Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountability.
  • Sufficient hands-on technical background to guide production troubleshooting across Linux-based systems and network services. You can interpret logs, metrics, alerts, and packet-level evidence and make sound operational decisions.
  • Working knowledge of networking concepts and technologies such as TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability. Deep expertise in every area is not required.
  • Experience leading or coordinating significant incidents and customer-impacting escalations in an always-on service environment.
  • Working knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring and alerting, incident management, and post-incident improvement.
  • An automation mindset and experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work and improve consistency.
  • Experience working with geographically distributed teams and cross-functional partners in software, platform, infrastructure, or cloud services.
  • Strong written and verbal communication skills, sound judgment under pressure, and consistent attention to detail.
  • Experience using or evaluating AI-assisted tools to improve analysis, decision-making, automation, or team workflows, with appropriate attention to accuracy, security, and operational risk.

Preferred qualifications

  • Experience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environment.
  • Familiarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public-cloud networking.
  • Experience improving observability, change safety, capacity management, or operational readiness for production services.
  • Experience with IT service management practices, including incident, change, and problem management.
  • Relevant certifications such as CCNA, CCNP, Azure/AWS/GCP related

Informations complémentaires

 

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity,  veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.  

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Skills Required

  • Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations
  • Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountability
  • Hands-on technical background guiding production troubleshooting across Linux-based systems and network services
  • Working knowledge of TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability
  • Experience leading or coordinating significant incidents and customer-impacting escalations in an always-on service environment
  • Working knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring, alerting, incident management, and post-incident improvement
  • Experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work
  • Experience working with geographically distributed teams and cross-functional partners
  • Strong written and verbal communication skills, sound judgment under pressure, and attention to detail
  • Experience using or evaluating AI-assisted tools for analysis, decision-making, automation, or team workflows
  • Experience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environment
  • Familiarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public-cloud networking
  • Experience improving observability, change safety, capacity management, or operational readiness for production services
  • Experience with IT service management practices, including incident, change, and problem management
  • Relevant certifications such as CCNA, CCNP, or Azure, AWS, or GCP certifications

What the Team is Saying

Shanequa
Katya
Suzanne
Alexander
Jaime
Pat
Brady
Hasan
Jamil
Viviana

ServiceNow Compensation & Benefits Highlights

  • Healthcare Strength Health coverage is described as comprehensive with multiple plan choices and strong perceived coverage, alongside mental-health resources and wellbeing support. Company materials and employer-verified summaries also note inclusive care options and supportive programs.
  • Parental & Family Support Parental leave and family-planning support are characterized as generous, with fully paid leave and resources such as fertility, caregiving, and adoption assistance. Backup care and other family-focused programs are also highlighted as part of the package.
  • Equity Value & Accessibility Equity components like RSUs and an employee stock purchase plan are presented as meaningful, widely available parts of total rewards. Many role and benefits overviews emphasize equity’s role in boosting overall compensation alongside bonuses.

ServiceNow Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
29,000 Employees
Year Founded: 2004

What We Do

As the AI platform for business transformation, we're putting AI to work across organizations — freeing people for work that matters. Making old tech work with new tech. Reaching across departments, from the front office to the back office and every office in between. Our ambition? To become the AI defining enterprise software company of the 21st century (or "AI DESCO21C," as we like to call it). With more than 8,400+ customers, we serve approximately 90% of the Fortune 500®, and we're proud to be a Fortune 100 Best Companies to Work For® and World's Most Admired Companies™. Explore your future career with us, visit www.careers.servicenow.com From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Why Work With Us

By joining ServiceNow, you are part of an ambitious team of change-makers who have a restless curiosity and a drive for ingenuity. We're committed to helping our people do their best work and live their best lives so we can fulfill our purpose together. At the fastest-growing enterprise software company, you can grow your career faster.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

ServiceNow Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

At ServiceNow, we lead with flexibility and trust. For some, home is the primary workplace. For those who come into a ServiceNow workplace, you are empowered to make team-guided and individual-led decisions on how and when you use the workplace.

Typical time on-site: Flexible
Company Office Image
HQSanta Clara, CA
Heredia
Ciudad de México
District of Columbia
Osaka
Aarhus, DK
Aarhus, DK
Company Office Image
Addison, TX
Amsterdam, North Holland
Atlanta, Georgia
Auckland, Auckland
Bangkok, Bangkok
Bengaluru, Karnataka
Bengaluru, Karnataka
Berlin, Berlin
Brasília, Federal District
Brisbane, Queensland
Brussels, BE
Cairo, Cairo Governorate
Canberra, Australian Capital Territory
Charlottesville, Virginia
Company Office Image
Chicago, IL
Deerfield, Illinois
Company Office Image
Denver, CO
Dubai, Dubai
Dublin, Leinster
Düsseldorf, Nordrhein-Westfalen
Frankfurt am Main, Hesse
Goteborg, Västra Götaland County
Gurugram, Haryana
Hamburg, Hamburg
Hanyang, Seoul
Helsinki, Uusimaa
Hong Kong, Hong Kong
Houston, TX
Hyderabad, Telangana
Issy-les-Moulineaux, Île-de-France
Johannesburg, Gauteng
Kirkland, WA
Lausanne, Vaud
Lille, Hauts de France
London, England
London, England
Madrid, Community of Madrid
Melbourne, Victoria
Milano, Lombardia
Milwaukee, WI
Company Office Image
Montréal, QC
Mumbai, Maharashtra
Munich, Bavaria
Company Office Image
New York, NY
Opfikon, Zürich
Orlando, FL
Oslo, Oslo
Perth, Western Australia
Petah Tikva, Central District
Company Office Image
Pleasanton, CA
Riyadh, Riyadh Province
Rome, Lazio
Company Office Image
San Diego, CA
San Francisco, Heredia
Company Office Image
San Francisco, CA
São Paulo, SP
Singapore, SG
Solna, Stockholm County
Sydney, New South Wales
Tokyo, Tokyo
Toronto, Ontario
Vancouver, British Columbia
Company Office Image
Vienna, VA
Vienna, AT
Company Office Image
Waltham, MA
Washington, DC
Wellington, Wellington
West Palm Beach, Florida
Learn more

Similar Jobs

ServiceNow Logo ServiceNow

Enterprise Account Exec

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Toronto, ON, CAN
29000 Employees
100K-175K Annually

ServiceNow Logo ServiceNow

Enterprise Account Executive

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Toronto, ON, CAN
29000 Employees

ServiceNow Logo ServiceNow

Database Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Toronto, ON, CAN
29000 Employees
109K-192K Annually

ServiceNow Logo ServiceNow

Dir, Sales, Cybersecurity Canada (Armis)

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Toronto, ON, CAN
29000 Employees
168K-200K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account