Principal Site Reliability Engineer

Posted One Month Ago
Be an Early Applicant
Buffalo, NY, USA
In-Office
140K-233K Annually
Senior level
Fintech
The Role
Lead design and implementation of highly available, fault-tolerant platform architectures. Define SLOs/SLAs, drive incident and problem management, build observability and automation, mentor engineers, and partner with stakeholders to improve reliability, performance, and compliance.
Summary Generated by Built In

Overview: 

Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, observability, and performance. Serves as a mentor and technical leader for less experienced engineers across Technology. 

Primary Responsibilities: 

• Accountable for defining and driving service reliability standards, including SLOs, SLAs, and error budgets across platforms. 
• Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements. 
• Lead incident management practices, including detection, response, escalation, and recovery processes. 
• Drive problem management and root cause analysis to prevent systemic issues. 
• Develop and promote observability strategies, including logging, monitoring, alerting, and tracing. 
• Lead automation initiatives for self-healing systems and operational workflows. 
• Contribute to and review technical roadmaps with reliability and performance considerations. 
• Partner with development, infrastructure, cybersecurity, and architecture teams. 
• Serve as a technical authority for performance, resilience, and capacity planning. 
• Drive production readiness practices including performance testing and failover capabilities. 
• Lead cross-team reliability improvement initiatives. 
• Participate in and lead post-incident reviews ensuring actionable outcomes. 
• Mentor engineers on reliability engineering and best practices. 
• Engage with stakeholders to identify risks and optimization opportunities. 
• Ensure adherence to risk and regulatory standards and escalate issues when needed. 
• Maintain internal control standards and compliance expectations. 

Scope of Responsibilities: 

Applies expert-level SRE practices across multiple platforms. Drives enterprise-wide reliability improvements and influences technical direction without direct authority. 

Supervisory/Managerial Responsibilities: 

No supervisory responsibilities. 

Education and Experience Required: 

Associate’s degree and a minimum of 9 years’ systems analysis and/ or application development work experience or Bachelor's degree and a minimum of 7 years' systems analysis and/ or application development work experience. In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/ or application development work experience.  

Expert experience in system design, reliability engineering, and production operations. 
Advanced proficiency in at least one programming or scripting language. 

Education and Experience Preferred: 

Experience with observability and incident management tooling. 
Experience with cloud platforms such as AWS or Azure. 
Strong understanding of CI/CD, DevOps, and SDLC practices. 
Experience defining and implementing SLO/SLI frameworks. 
Experience in regulated environments such as financial services. 
Strong communication and stakeholder management skills. 

 

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

LocationBuffalo, New York, United States of America

Skills Required

  • Bachelor's degree plus minimum 7 years systems analysis and/or application development experience (or Associate's + 9 years; or 11 years combined education/work in lieu of degree)
  • Expert experience in system design, reliability engineering, and production operations
  • Advanced proficiency in at least one programming or scripting language
  • Experience defining and implementing SLO/SLI frameworks
  • Experience with observability and incident management tooling
  • Experience with cloud platforms such as AWS or Azure
  • Strong understanding of CI/CD, DevOps, and SDLC practices
  • Experience in regulated environments such as financial services
  • Strong communication and stakeholder management skills

M&T Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about M&T Bank and has not been reviewed or approved by M&T Bank.

  • Retirement Support — Retirement benefits are positioned as a strong pillar, including a 401(k) match and the possibility of an additional employer contribution, plus access to an employee stock purchase plan.
  • Leave & Time Off Breadth — Time-off offerings are framed as competitive, with a flexible PTO approach and paid volunteer time called out as a meaningful add-on to standard leave.
  • Wellbeing & Lifestyle Benefits — Wellbeing support appears comparatively robust, highlighted by mental-health therapy/coaching sessions and broader wellness programming alongside community-oriented perks.

M&T Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Buffalo, NY
21,590 Employees
Year Founded: 1856

What We Do

M&T Bank is a multi-state community-focused bank serving New York, Maryland, New Jersey, Pennsylvania, Delaware, Connecticut, Virginia, West Virginia and Washington, D.C. Founded in 1856, the company provides banking, investment, insurance and mortgage financial services to more than 3.6 million consumer, business and government clients.

Similar Jobs

Factset Logo Factset

Site Reliability Engineer

Aerospace • Big Data • Fintech • Software • Analytics
Hybrid
2 Locations
10310 Employees
190K-220K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
169K-305K Annually

CoreWeave Logo CoreWeave

Senior Software Engineer

Cloud • Information Technology • Machine Learning
In-Office
7 Locations
1450 Employees
153K-204K Annually
In-Office or Remote
New York, NY, USA
20 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account