Principal Site Reliability Engineer

Posted One Month Ago
Be an Early Applicant
Buffalo, NY, USA
In-Office
140K-233K Annually
Senior level
Fintech
The Role
Lead design and implementation of highly available, fault-tolerant platform architectures. Define SLOs/SLAs, drive incident and problem management, build observability and automation, mentor engineers, and partner with stakeholders to improve reliability, performance, and compliance.
Summary Generated by Built In

Overview: 

Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, observability, and performance. Serves as a mentor and technical leader for less experienced engineers across Technology. 

Primary Responsibilities: 

• Accountable for defining and driving service reliability standards, including SLOs, SLAs, and error budgets across platforms. 
• Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements. 
• Lead incident management practices, including detection, response, escalation, and recovery processes. 
• Drive problem management and root cause analysis to prevent systemic issues. 
• Develop and promote observability strategies, including logging, monitoring, alerting, and tracing. 
• Lead automation initiatives for self-healing systems and operational workflows. 
• Contribute to and review technical roadmaps with reliability and performance considerations. 
• Partner with development, infrastructure, cybersecurity, and architecture teams. 
• Serve as a technical authority for performance, resilience, and capacity planning. 
• Drive production readiness practices including performance testing and failover capabilities. 
• Lead cross-team reliability improvement initiatives. 
• Participate in and lead post-incident reviews ensuring actionable outcomes. 
• Mentor engineers on reliability engineering and best practices. 
• Engage with stakeholders to identify risks and optimization opportunities. 
• Ensure adherence to risk and regulatory standards and escalate issues when needed. 
• Maintain internal control standards and compliance expectations. 

Scope of Responsibilities: 

Applies expert-level SRE practices across multiple platforms. Drives enterprise-wide reliability improvements and influences technical direction without direct authority. 

Supervisory/Managerial Responsibilities: 

No supervisory responsibilities. 

Education and Experience Required: 

Associate’s degree and a minimum of 9 years’ systems analysis and/ or application development work experience or Bachelor's degree and a minimum of 7 years' systems analysis and/ or application development work experience. In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/ or application development work experience.  

Expert experience in system design, reliability engineering, and production operations. 
Advanced proficiency in at least one programming or scripting language. 

Education and Experience Preferred: 

Experience with observability and incident management tooling. 
Experience with cloud platforms such as AWS or Azure. 
Strong understanding of CI/CD, DevOps, and SDLC practices. 
Experience defining and implementing SLO/SLI frameworks. 
Experience in regulated environments such as financial services. 
Strong communication and stakeholder management skills. 

 

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

LocationBuffalo, New York, United States of America

Skills Required

  • Bachelor's degree plus minimum 7 years systems analysis and/or application development experience (or Associate's + 9 years; or 11 years combined education/work in lieu of degree)
  • Expert experience in system design, reliability engineering, and production operations
  • Advanced proficiency in at least one programming or scripting language
  • Experience defining and implementing SLO/SLI frameworks
  • Experience with observability and incident management tooling
  • Experience with cloud platforms such as AWS or Azure
  • Strong understanding of CI/CD, DevOps, and SDLC practices
  • Experience in regulated environments such as financial services
  • Strong communication and stakeholder management skills

M&T Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about M&T Bank and has not been reviewed or approved by M&T Bank.

  • Retirement Support Retirement benefits are positioned as a strong pillar, including a 401(k) match and the possibility of an additional employer contribution, plus access to an employee stock purchase plan.
  • Leave & Time Off Breadth Time-off offerings are framed as competitive, with a flexible PTO approach and paid volunteer time called out as a meaningful add-on to standard leave.
  • Wellbeing & Lifestyle Benefits Wellbeing support appears comparatively robust, highlighted by mental-health therapy/coaching sessions and broader wellness programming alongside community-oriented perks.

M&T Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Buffalo, NY
21,590 Employees
Year Founded: 1856

What We Do

M&T Bank is a multi-state community-focused bank serving New York, Maryland, New Jersey, Pennsylvania, Delaware, Connecticut, Virginia, West Virginia and Washington, D.C. Founded in 1856, the company provides banking, investment, insurance and mortgage financial services to more than 3.6 million consumer, business and government clients.

Similar Jobs

Capco Logo Capco

Site Reliability Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Hybrid
New York, NY, USA
6000 Employees
146K-183K Annually

TIDAL Logo TIDAL

Creative Director

Consumer Web • Information Technology • Mobile • Music • News + Entertainment • Software
Remote or Hybrid
New York, NY, USA
450 Employees
252K-377K Annually
Hybrid
4 Locations
289097 Employees

Canoe Logo Canoe

Data Analyst

Artificial Intelligence • Fintech • Information Technology • Machine Learning • Financial Services
Hybrid
New York City, NY, USA
180 Employees
70K-100K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account