Site Reliability Engineer IV

Reposted 3 Days Ago
Buffalo, NY, USA
In-Office
140K-233K Annually
Senior level
Fintech
The Role
Lead design and implementation of highly available, fault-tolerant platform architectures. Define SLO/SLI/SLA frameworks, drive observability and monitoring (logs, tracing, Dynatrace, OTel), lead incident and problem management, automate self-healing and deployment workflows, develop IaC (Terraform), optimize Azure/cloud environments, and mentor engineers to improve reliability, performance, and operational maturity across the enterprise.
Summary Generated by Built In
Overview

Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, observability, testing, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, resiliency, performance, and operational maturity. Serves as a mentor and technical leader for less experienced engineers across Technology.

Primary Responsibilities
  • Accountable for defining and driving service reliability standards, including SLOs, SLAs, SLIs, and error budgets across platforms.
  • Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements.
  • Lead initiatives to improve system reliability, availability, performance, and operational excellence through automation and engineering best practices.
  • Develop and promote observability strategies leveraging logging, monitoring, alerting, distributed tracing, OpenTelemetry (OTel), Dynatrace, dashboards, and telemetry analytics.
  • Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
  • Lead incident management practices, including detection, response, escalation, recovery, and coordination of high-severity production events.
  • Drive problem management and Root Cause Analysis (RCA) activities to prevent systemic issues and ensure corrective actions are implemented.
  • Lead automation initiatives for self-healing systems, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to build reliable, observable, and scalable services throughout the Software Development Lifecycle (SDLC).
  • Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
  • Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for infrastructure provisioning, configuration management, and environment standardization.
  • Support and optimize cloud environments, including Microsoft Azure services, deployment automation, scaling strategies, and application lifecycle management.
  • Utilize cloud-native monitoring and operational tools to improve platform visibility, reliability, and performance.
  • Serve as a technical authority for performance engineering, resilience, capacity planning, and workload optimization.
  • Drive production readiness practices, including performance testing, resiliency testing, failover validation, disaster recovery preparedness, and operational readiness reviews.
  • Review architectural designs and technical roadmaps, providing recommendations to improve reliability, scalability, resiliency, and operational efficiency.
  • Lead cross-team reliability improvement initiatives and influence enterprise engineering standards.
  • Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
  • Partner with development, infrastructure, cybersecurity, architecture, and support teams to identify risks, drive continuous improvement, and optimize platform performance.
  • Participate in and lead post-incident reviews, ensuring actionable outcomes and measurable improvements.
  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
  • Present reliability initiatives, operational metrics, and engineering recommendations in architecture reviews, technical forums, and leadership discussions.
  • Mentor engineers on reliability engineering, observability, cloud engineering, automation, and operational best practices.
  • Engage with stakeholders to identify risks, dependencies, and optimization opportunities.
  • Ensure adherence to risk and regulatory standards and escalate issues when needed.
  • Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
  • Maintain internal control standards, including timely implementation of audit findings, regulatory requirements, and compliance expectations.
  • Complete other related duties as assigned.
Scope of Responsibilities

Applies expert-level Site Reliability Engineering practices across multiple platforms. Drives enterprise-wide reliability improvements and influences technical direction without direct authority. Provides technical leadership, mentorship, and guidance to engineers and project teams.

Supervisory/Managerial Responsibilities

No supervisory responsibilities.

Education and Experience Required

Associate’s degree and a minimum of 9 years’ systems analysis and/or application development work experience or Bachelor’s degree and a minimum of 7 years’ systems analysis and/or application development work experience. In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/or application development work experience.

Expert experience in system design, reliability engineering, and production operations.

Advanced proficiency in at least one programming or scripting language.

Education and Experience Preferred
  • Experience with observability and incident management tooling.
  • Experience with cloud platforms such as AWS or Azure.
  • Strong understanding of CI/CD, DevOps, and SDLC practices.
  • Experience defining and implementing SLO/SLI frameworks.
  • Experience in regulated environments such as financial services.
  • Strong communication and stakeholder management skills.
  • Experience with Infrastructure as Code (IaC), including Terraform.
  • Experience with monitoring and observability platforms such as Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, or similar tools.
  • Experience with automated testing, deployment automation, and reliability engineering practices.
  • Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures.
  • Experience working in Agile and DevOps operating models.
  • Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.
  • Industry certifications related to Cloud Engineering, Azure, AWS, Terraform, or Site Reliability Engineering preferred.

#LI-JB3

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

LocationBuffalo, New York, United States of America

Skills Required

  • Associate's degree and minimum 9 years systems analysis/application development experience OR Bachelor's degree and minimum 7 years systems analysis/application development experience (or equivalent combination).
  • Expert experience in system design, reliability engineering, and production operations.
  • Advanced proficiency in at least one programming or scripting language.

M&T Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about M&T Bank and has not been reviewed or approved by M&T Bank.

  • Retirement Support Retirement benefits are positioned as a strong pillar, including a 401(k) match and the possibility of an additional employer contribution, plus access to an employee stock purchase plan.
  • Leave & Time Off Breadth Time-off offerings are framed as competitive, with a flexible PTO approach and paid volunteer time called out as a meaningful add-on to standard leave.
  • Wellbeing & Lifestyle Benefits Wellbeing support appears comparatively robust, highlighted by mental-health therapy/coaching sessions and broader wellness programming alongside community-oriented perks.

M&T Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Buffalo, NY
21,590 Employees
Year Founded: 1856

What We Do

M&T Bank is a multi-state community-focused bank serving New York, Maryland, New Jersey, Pennsylvania, Delaware, Connecticut, Virginia, West Virginia and Washington, D.C. Founded in 1856, the company provides banking, investment, insurance and mortgage financial services to more than 3.6 million consumer, business and government clients.

Similar Jobs

The Walt Disney Company Logo The Walt Disney Company

Site Reliability Engineer

Digital Media • Gaming • News + Entertainment • Sports
In-Office
New York, NY, USA
219548 Employees
123K-165K Annually

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
New York City, NY, USA
5550 Employees
111K-218K Annually

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
10 Locations
5550 Employees
127K-249K Annually

ASAPP Logo ASAPP

Senior Site Reliability Engineer

Artificial Intelligence • Machine Learning • Natural Language Processing • Software
Hybrid
2 Locations
389 Employees
150K-175K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account