GenAI Platform Engineering Lead

Posted 9 Days Ago
Be an Early Applicant
Buffalo, NY, USA
In-Office
116K-194K Annually
Expert/Leader
Fintech
The Role
Leads engineering teams responsible for the reliability, observability, security, cost efficiency, and operational readiness of a bank’s AI platform. Oversees incident response, runbooks, monitoring, infrastructure automation, production support, Azure services, Terraform-based deployments, AI-specific operational controls, risk compliance, budgets, vendor relationships, and client communications. Provides hands-on technical direction during complex production issues while managing staffing, performance, projects, and organizational objectives.
Summary Generated by Built In

Manages the activities of Engineering Team Leaders, Engineering Supervisors and/or engineering units responsible for the reliability, observability and operational readiness of the Bank’s AI platform. Provides day-to-day direction for the teams and applications in alignment with departmental goals and the needs of the clients they support.

Serves as the technical lead for AI operations and service assurance, with accountability for the operational layer of the AI platform. Works closely with the Platform Engineering Manager, who owns the broader platform, to ensure production systems are secure, resilient, observable, cost-effective and operationally ready. Oversees incident response practices, runbooks, monitoring, infrastructure automation and production support.

Responsible for managing client relationships and expectations, prioritizing the project queue and achieving individual and organizational objectives at minimum cost.

Primary Responsibilities

  • Lead the reliability, observability and operational readiness of the AI platform, including production monitoring, incident response, service assurance, infrastructure automation and operational controls.
  • Establish and maintain operational runbooks, escalation procedures, service-level indicators, service-level objectives and incident response practices.
  • Measure operational performance through platform availability, incident response time, mean time to recovery, model latency, token consumption and infrastructure cost.
  • Provide leadership during production incidents. Coordinate troubleshooting, communication, escalation, root-cause analysis and corrective actions.
  • Build and maintain observability pipelines using OpenTelemetry, Prometheus, Grafana, Azure Monitor, Log Analytics and/or comparable technologies.
  • Oversee the deployment and operation of Azure infrastructure and services, including Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID and infrastructure managed through Terraform.
  • Promote infrastructure-as-code and automated deployment practices using Terraform, GitHub Actions, GitLab CI and/or comparable tools. Ensure platform changes are deployed through controlled pipelines rather than manual processes.
  • Drive automation of recurring operational activities using Python, Bash, PowerShell and other appropriate scripting technologies.
  • Oversee AI-specific operational capabilities, which may include token cost tracking, model latency monitoring, provider failover, caller-level rate limiting, prompt logging, personally identifiable information interception and infrastructure-layer content filtering.
  • Partner with cybersecurity and risk teams to support Security Information and Event Management integrations and security event feeds from application infrastructure.
  • Manage and participate in consultations with client management to analyze short-range business requirements and recommend innovations that anticipate the future impact of changing business and technology needs. Build and maintain positive client relationships.
  • Monitor technology direction, industry trends and vendor applications related to AI platforms, site reliability engineering, cloud infrastructure, observability and service assurance.
  • Research and initiate changes to existing processes, technologies and operating models when necessary.
  • Lead vendor and product analysis and provide recommendations.
  • Oversee application development support, testing efforts, technology infrastructure, project management and other assigned technology domains.
  • Serve as a subject matter expert for AI platform operations, production reliability, observability and service assurance.
  • Build rapport across the organization and maintain a professional level of communication and cooperation with technology, business, risk, cybersecurity and vendor partners.
  • Maintain relationships with vendors and professional organizations.
  • Direct team activities, assign personnel to projects and provide technical and operational guidance.
  • Ensure schedules and commitments are completed. Lead short-term staffing and capacity planning.
  • Implement technology consistent with Division standards and long-range plans. Ensure adherence to Department and Technology standards and procedures, including documentation, audit trail and change management requirements.
  • Translate business and operational requirements as needed to assist staff in preparing detailed specifications for system enhancements.
  • Evaluate and manage recommended designs based on business, operational and technology requirements. Identify, communicate and resolve issues and concerns.
  • Manage project plans and coordinate major project and production-readiness activities. Remain current on work outside the team that may affect the team, platform or client environment.
  • Develop and manage multiple cost center budgets, including cloud infrastructure and AI platform operating costs.
  • Recommend and implement policies and procedures that improve the performance, reliability and effectiveness of the Department.
  • Exercise the usual authority of a manager concerning staffing, performance appraisals, promotions, salary recommendations, performance management and terminations.
  • Understand and adhere to the Company’s risk and regulatory standards, policies and controls in accordance with the Company’s Risk Appetite. Design, implement, maintain and enhance internal controls to mitigate risk on an ongoing basis. Identify risk-related issues requiring escalation to management.
  • Promote an environment that supports belonging and reflects the M&T Bank brand.
  • Maintain M&T internal control standards, including the timely implementation of internal and external audit points and resolution of issues raised by external regulators, as applicable.
  • Complete other related duties as assigned.

Scope of Responsibilities

Oversees a team where the majority of employees are engineers, architect individual contributors, Engineering Supervisors and/or Engineering Team Leaders. Leads the operational capabilities supporting the AI platform and partners closely with the Platform Engineering Manager, application teams, cybersecurity, risk, architecture and other technology stakeholders.

This role requires both people leadership and technical depth. The manager is expected to provide hands-on technical direction, support complex production troubleshooting and ensure the team’s operational practices meet the Bank’s reliability, security, risk and regulatory expectations.

Supervisory/Managerial Responsibilities

5 to 10

Skills and Success Indicators

  • Operational ownership: Measures success through platform availability, incident response time, mean time to recovery, operational readiness and the quality of documented runbooks.
  • Technical depth: Can review telemetry, interpret traces, troubleshoot infrastructure and gateway configurations and provide technical direction during complex production issues.
  • Leadership: Develops engineers and leaders, establishes clear accountability and creates an environment that supports collaboration, belonging and continuous improvement.
  • Automation mindset: Identifies repeatable operational activities and moves them toward scripted, pipeline-based and self-service solutions.
  • Cost awareness: Treats token consumption, infrastructure usage and cloud expense as core operational metrics.
  • Risk and control orientation: Builds auditability, security, change management and regulatory expectations into operational processes.
  • Client partnership: Builds trust with business and technology stakeholders through clear communication, dependable execution and transparent incident management.

Education and Experience Required

  • A combined minimum of 9 years’ higher education and/or work experience, including a minimum of 4 years’ engineering and/or architecture experience and 3 years leadership experience
  • Minimum of 5 years’ experience in site reliability engineering, platform operations, infrastructure engineering or a related discipline, including experience supporting production systems through an on-call model
  • Hands-on experience building, operating and troubleshooting Azure infrastructure and services, including Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID and Terraform
  • Experience building and supporting observability pipelines using OpenTelemetry, Prometheus, Grafana, Azure-native technologies and/or comparable tools
  • Strong understanding of metrics, traces and logs, including when and how each should be used to monitor and troubleshoot production systems
  • Experience implementing continuous integration and continuous delivery practices for infrastructure using Terraform, GitHub Actions, GitLab CI, infrastructure-as-code patterns and/or comparable technologies
  • Strong scripting and automation skills using Python, Bash, PowerShell and/or comparable languages
  • Experience leading or supporting production incident response, including troubleshooting, escalation, root-cause analysis and service restoration
  • Capable of working on multiple projects of a complex nature
  • Proficiency with project management, word processing and spreadsheet applications
  • Complete understanding of the system development life cycle
  • Excellent problem-solving skills to assist in issue resolution
  • Familiarity with application development software, cloud infrastructure and hardware platforms
  • Excellent verbal and written communication skills
  • Excellent analytical and decision-making skills
  • Strong project management and presentation skills
  • Experience encouraging teamwork and serving as a role model when leading and directing others
  • Understanding of the technical, business, operational, risk and cost impacts of a project, platform or production issue

Education and Experience Preferred

  • Bachelor’s degree
  • Minimum of 10 years’ technology management, site reliability engineering, platform operations or large program leadership experience
  • Experience operating AI or large language model platforms in a production environment
  • Experience with large language model operational concerns, including token cost tracking, model latency monitoring, provider failover and rate limiting by caller identity
  • Experience integrating application infrastructure with Security Information and Event Management platforms and security event feeds
  • Familiarity with AI gateway patterns, including prompt logging, personally identifiable information interception and infrastructure-layer content filtering
  • Experience working in a regulated environment where audit trails, access controls and change management practices are required
  • Extensive application and product knowledge within AI platforms, cloud infrastructure, site reliability engineering, observability and/or service assurance
  • Subject matter expert understanding of supported applications, with advanced knowledge of interfacing and integrated applications
  • Understanding of multiple business areas and their functions
  • Proven mentoring and leadership capabilities
  • Experience with the technologies, applications and functions of the area being led
  • Good understanding of the Bank’s application framework
  • Awareness of the Bank’s business plan and strategic objectives, with the ability to help shape direction
  • Self-motivated with the ability to motivate and develop others
  • Understanding of the supported businesses and their terminology

Skills and Success Indicators

  • Operational ownership: Measures success through platform availability, incident response time, mean time to recovery, operational readiness and the quality of documented runbooks.
  • Technical depth: Can review telemetry, interpret traces, troubleshoot infrastructure and gateway configurations and provide technical direction during complex production issues.
  • Leadership: Develops engineers and leaders, establishes clear accountability and creates an environment that supports collaboration, belonging and continuous improvement.
  • Automation mindset: Identifies repeatable operational activities and moves them toward scripted, pipeline-based and self-service solutions.
  • Cost awareness: Treats token consumption, infrastructure usage and cloud expense as core operational metrics.
  • Risk and control orientation: Builds auditability, security, change management and regulatory expectations into operational processes.
  • Client partnership: Builds trust with business and technology stakeholders through clear communication, dependable execution and transparent incident management.

#LI-JB3

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $116,400.00 - $194,000.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

LocationBuffalo, New York, United States of America

Skills Required

  • A combined minimum of 9 years of higher education and/or work experience
  • At least 4 years of engineering and/or architecture experience
  • At least 3 years of leadership experience
  • At least 5 years of site reliability engineering, platform operations, infrastructure engineering, or related experience
  • Experience supporting production systems through an on-call model
  • Hands-on experience building, operating, and troubleshooting Azure infrastructure and services, including Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID, and Terraform
  • Experience building and supporting observability pipelines using OpenTelemetry, Prometheus, Grafana, Azure-native technologies, or comparable tools
  • Strong understanding of metrics, traces, and logs for production monitoring and troubleshooting
  • Experience implementing infrastructure CI/CD using Terraform, GitHub Actions, GitLab CI, infrastructure-as-code patterns, or comparable technologies
  • Strong scripting and automation skills using Python, Bash, PowerShell, or comparable languages
  • Experience leading or supporting production incident response, troubleshooting, escalation, root-cause analysis, and service restoration
  • Ability to work on multiple complex projects
  • Proficiency with project management, word processing, and spreadsheet applications
  • Complete understanding of the system development life cycle
  • Excellent problem-solving, verbal and written communication, analytical, and decision-making skills
  • Strong project management and presentation skills
  • Experience encouraging teamwork and leading or directing others
  • Understanding of technical, business, operational, risk, and cost impacts of projects, platforms, or production issues
  • Bachelor’s degree
  • At least 10 years of technology management, site reliability engineering, platform operations, or large program leadership experience
  • Experience operating AI or large language model platforms in production
  • Experience with LLM token cost tracking, model latency monitoring, provider failover, and caller-based rate limiting
  • Experience integrating application infrastructure with SIEM platforms and security event feeds
  • Familiarity with AI gateway patterns, prompt logging, PII interception, and infrastructure-layer content filtering
  • Experience working in a regulated environment with audit trails, access controls, and change management
  • Extensive knowledge of AI platforms, cloud infrastructure, SRE, observability, and/or service assurance
  • Proven mentoring and leadership capabilities

M&T Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about M&T Bank and has not been reviewed or approved by M&T Bank.

  • Retirement Support Retirement benefits are positioned as a strong pillar, including a 401(k) match and the possibility of an additional employer contribution, plus access to an employee stock purchase plan.
  • Leave & Time Off Breadth Time-off offerings are framed as competitive, with a flexible PTO approach and paid volunteer time called out as a meaningful add-on to standard leave.
  • Wellbeing & Lifestyle Benefits Wellbeing support appears comparatively robust, highlighted by mental-health therapy/coaching sessions and broader wellness programming alongside community-oriented perks.

M&T Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Buffalo, NY
21,590 Employees
Year Founded: 1856

What We Do

M&T Bank is a multi-state community-focused bank serving New York, Maryland, New Jersey, Pennsylvania, Delaware, Connecticut, Virginia, West Virginia and Washington, D.C. Founded in 1856, the company provides banking, investment, insurance and mortgage financial services to more than 3.6 million consumer, business and government clients.

Similar Jobs

CrowdStrike Logo CrowdStrike

Infrastructure Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
100K-155K Annually

CrowdStrike Logo CrowdStrike

Senior Platform Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

Enverus Logo Enverus

Project Manager

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
2 Locations
1800 Employees

Enverus Logo Enverus

Director Of Sales

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
5 Locations
1800 Employees
140K-180K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account