Site Reliability Engineering (SRE) Manager

Posted 9 Days Ago
Be an Early Applicant
Buffalo, NY, USA
In-Office
140K-233K Annually
Expert/Leader
Fintech
The Role
Leads SRE, production support, observability, incident management, automation, cloud reliability, and platform operations teams. Defines reliability strategies, SLIs/SLOs, monitoring standards, disaster recovery, incident response, root cause analysis, and operational governance. Drives Azure cloud modernization, infrastructure automation, AI-assisted operations, and continuous improvement while managing and developing engineering staff responsible for critical business systems.
Summary Generated by Built In
This role is four days onsite at either Seneca One Buffalo, NY location or Wilmington Center Wilmington, DE, location with the flexibility to work from home one day per weekOverview

The Site Reliability Engineering (SRE) Manager leads teams responsible for the reliability, availability, performance, and operational excellence of critical business applications and platforms. This role combines engineering leadership with deep expertise in production operations, observability, automation, incident management, and cloud technologies.

 

The SRE Manager partners with Engineering, Architecture, Infrastructure, Security, Product, and Business stakeholders to ensure systems are resilient, scalable, secure, and supportable. The role is accountable for driving operational excellence through automation, reliability engineering practices, and continuous improvement while developing high-performing SRE and Production Support teams.

 

Primary Responsibilities

Reliability & Operational Excellence

  • Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.
  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.
  • Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.
  • Drive continuous improvement of application stability, service availability, and customer experience.

Incident & Problem Management

  • Lead major incident response and escalation management for critical production issues.
  • Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.
  • Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.
  • Provide executive-level communication during significant incidents and service disruptions.

Observability & Automation

  • Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.
  • Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.
  • Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.
  • Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.

Cloud & Platform Reliability

  • Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.
  • Ensure applications are designed and operated using resilient, scalable, and supportable architectures.
  • Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.
  • Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.

AI & Modern Operations

  • Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.
  • Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.
  • Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.

People Leadership

  • Recruit, develop, coach, and retain high-performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.
  • Establish career paths, skill development plans, and succession strategies.
  • Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.
  • Manage staffing, performance management, compensation recommendations, and organizational development activities.

Risk & Governance

  • Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.
  • Identify and escalate operational risks impacting critical services or customer experiences.
  • Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.

 

Scope of Responsibilities

Leads teams responsible for:

  • Site Reliability Engineering (SRE)
  • Production Support
  • Observability Engineering
  • Incident Management
  • Operational Automation
  • Cloud Reliability
  • Platform Operations

Responsible for reliability and operational health across multiple applications, platforms, cloud services, and vendor-supported solutions.

 

Supervisory Responsibilities

Typically manages 10-20 direct and indirect reports including SRE Engineers, Production Engineers, Technical Leads, and Engineering Managers.

 

Education & Experience Required
  • 10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities.
  • 5+ years of leadership experience managing engineering, operations, or SRE teams.
  • Experience managing production systems supporting critical business functions.
  • Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence.
  • Experience leading major incident response, root cause analysis, and service restoration efforts.
  • Experience with cloud platforms, distributed systems, APIs, and modern application architectures.
  • Strong communication, analytical, decision-making, and stakeholder management skills.

 

Preferred Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
  • Experience leading SRE or Production Engineering organizations.
  • Experience with Azure cloud technologies and cloud-native architectures.
  • Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
  • Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs.
  • Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices.
  • Experience implementing operational AI use cases including incident analysis, observability analytics, and automated diagnostics.
  • Financial services or other highly regulated industry experience preferred.

 

What Great Looks Like

A successful SRE Manager at M&T:

  • Delivers highly available and resilient customer-facing platforms.
  • Uses automation to eliminate operational toil and improve efficiency.
  • Reduces mean time to detect (MTTD) and mean time to restore (MTTR).
  • Establishes strong observability and operational intelligence capabilities.
  • Builds a culture of reliability, accountability, and continuous improvement.
  • Successfully integrates AI-assisted operations and automation into support workflows.
  • develops high-performing teams that balance reliability, speed, risk management, and customer experience.

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

LocationBuffalo, New York, United States of America

Skills Required

  • 10+ years of technology experience involving application support, infrastructure, cloud, software engineering, or reliability engineering
  • 5+ years of leadership experience managing engineering, operations, or SRE teams
  • Experience managing production systems supporting critical business functions
  • Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence
  • Experience leading major incident response, root cause analysis, and service restoration efforts
  • Experience with cloud platforms, distributed systems, APIs, and modern application architectures
  • Strong communication, analytical, decision-making, and stakeholder management skills
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field
  • Experience leading SRE or Production Engineering organizations
  • Experience with Azure cloud technologies and cloud-native architectures
  • Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry
  • Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs
  • Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices
  • Experience implementing operational AI use cases, including incident analysis, observability analytics, and automated diagnostics
  • Financial services or other highly regulated industry experience

M&T Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about M&T Bank and has not been reviewed or approved by M&T Bank.

  • Retirement Support — Retirement benefits are positioned as a strong pillar, including a 401(k) match and the possibility of an additional employer contribution, plus access to an employee stock purchase plan.
  • Leave & Time Off Breadth — Time-off offerings are framed as competitive, with a flexible PTO approach and paid volunteer time called out as a meaningful add-on to standard leave.
  • Wellbeing & Lifestyle Benefits — Wellbeing support appears comparatively robust, highlighted by mental-health therapy/coaching sessions and broader wellness programming alongside community-oriented perks.

M&T Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Buffalo, NY
21,590 Employees
Year Founded: 1856

What We Do

M&T Bank is a multi-state community-focused bank serving New York, Maryland, New Jersey, Pennsylvania, Delaware, Connecticut, Virginia, West Virginia and Washington, D.C. Founded in 1856, the company provides banking, investment, insurance and mortgage financial services to more than 3.6 million consumer, business and government clients.

Similar Jobs

T-Mobile Logo T-Mobile

Senior Data Engineer

Other • Utilities
In-Office
2 Locations
89016 Employees
105K-218K Annually

Comcast Logo Comcast

Analyst, Deal Desk

Digital Media • Information Technology • News + Entertainment
Hybrid
New York, NY, USA
115000 Employees
72K-108K Annually

PwC Logo PwC

New York - Tax LLM - Associate - Summer/Fall 2027

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
New York, NY, USA
370000 Employees
54K-142K Annually

PwC Logo PwC

AWS Cloud Infrastructure and Operations Delivery Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
67 Locations
370000 Employees
99K-232K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account