Site Reliability Engineer Lead

Posted 17 Days Ago
Be an Early Applicant
3 Locations
In-Office
Expert/Leader
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
The Role
Lead and build an SRE team to ensure reliability, scalability, and performance of infrastructure automation platforms. Define SRE frameworks (SLIs/SLOs/SLAs), drive observability and automation, manage incident response and RCAs, implement capacity planning, integrate tooling with CI/CD and ITSM, mentor engineers, and partner with stakeholders for operational excellence.
Summary Generated by Built In

Job Description:

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates’ physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve.
Bank of America is committed to an in-office culture with specific requirements for office-based attendance and which allows for an appropriate level of flexibility for our teammates and businesses based on role-specific considerations.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!
 

Job Description:

This job is responsible for building and leading a team to deliver technology products and services that meet business outcomes. Key responsibilities include developing a technology strategy, ensuring technology solutions comply with applicable standards, promoting design, engineering, and organizational practices, and advocating and advancing modern, Agile solution delivery practices. Job expectations may include coaching, mentoring, providing feedback and hands on career development, identifying emerging talent, fostering leadership skills, and managing stakeholders.

Overview:

Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the reliability, scalability, and performance of critical Infrastructure Automation platforms. This role will lead the design and implementation of SRE practices across a federated technology ecosystem, ensuring operational excellence through automation, observability, and resilient architecture.

The ideal candidate will bring deep expertise in distributed systems, cloud-native infrastructure, SaaS application support and DevOps/SRE principles, along with strong leadership and collaboration skills to influence cross-functional engineering and Production management teams and drive continuous improvement in service reliability.

Responsibilities:

SRE Strategy & Governance:

  • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols.
  • Establish governance models for reliability engineering across distributed teams.
  • Champion a culture of observability, proactive monitoring, and continuous feedback loops.

Reactive & Proactive Problem Management:

  • Lead root cause analysis (RCA) and post-incident reviews to identify systemic issues and prevent recurrence.
  • Implement proactive problem detection using telemetry, anomaly detection, and trend analysis.
  • Collaborate with engineering and operations teams to eliminate toil and reduce incident frequency and impact.

Capacity & Performance Management:

  • Develop and maintain capacity models to ensure systems scale efficiently with business demand.
  • Monitor performance trends and lead optimization efforts across infrastructure and applications.
  • Partner with finance and engineering teams to align capacity planning with cost and growth objectives.

Platform Reliability & Automation:

  • Drive automation of operational tasks including deployments, scaling, and recovery.
  • Integrate reliability tooling with CI/CD pipelines, ITSM platforms (e.g., ServiceNow), and observability systems.

Incident Management & Operational Excellence:

  • Oversee major incident response, escalation, and communication processes.
  • Develop and maintain runbooks, playbooks, and escalation protocols.
  • Drive continuous improvement through blameless retrospectives and operational reviews.

Technical Leadership:

  • Serve as a senior technical advisor and thought leader in SRE and platform engineering.
  • Mentor and guide SRE teams and partner with engineering leaders across the enterprise.
  • Provide input on staffing, tooling strategy, and budget planning for reliability initiatives.

Managerial Responsibilities:
This position may also have responsibilities for managing associates. At Bank of America, all managers at this level demonstrate the following responsibilities, in addition to those specific to the role, listed above.

  • Opportunity & Inclusion Champion: Models an inclusive environment for employees and clients, aligned to company Great Place to Work goals.
  • Manager of Process & Data: Demonstrates deep process knowledge, operational excellence and innovation through a focus on simplicity, data based decision making and continuous improvement.
  • Enterprise Advocate & Communicator: Communicates enterprise decisions, purpose, and results, and connects to team strategy, priorities and contributions.
  • Risk Manager: Ensures proper risk discipline, controls and culture are in place to identify, escalate and debate issues.
  • People Manager & Coach: Provides inspection, coaching and feedback to motivate, differentiate and improve performance.
  • Financial Steward: Actively manages expenses and budgets in alignment with objectives, making sound financial decisions.
  • Enterprise Talent Leader: Assesses talent and builds bench strength for roles across the organization.
  • Driver of Business Outcomes: Delivers results by effectively prioritizing, inspecting and appropriately delegating team work.

Required Qualifications:

  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments.
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing.
  • Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry).
  • Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python).
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes).
  • Demonstrated success in leading incident response and driving systemic improvements.
  • Experience with capacity planning, performance tuning, and cost optimization.
  • Excellent communication and stakeholder management skills, including executive engagement.

Desired Qualifications:

  • Experience with ITIL/ITSM processes and integration with platforms like ServiceNow.
  • Familiarity with security and compliance in regulated industries (e.g., financial services).
  • Background in performance engineering and infrastructure analytics.
  • Experience developing dashboards and metrics for operational health and reliability.

Skills:

  • Influence
  • Risk Management
  • Solution Design
  • Stakeholder Management
  • Technical Strategy Development
  • Analytical Thinking
  • Application Development
  • Collaboration
  • Result Orientation
  • Solution Delivery Process
  • Agile Practices
  • Architecture
  • Automation
  • Data Management
  • DevOps Practices

Shift:

1st shift (United States of America)

Hours Per Week: 

40

Skills Required

  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments
  • Deep understanding of Linux/Unix and Windows systems, networking, and distributed computing
  • Proven experience with observability stacks (Dynatrace, Grafana, Splunk, OpenTelemetry)
  • Expertise in infrastructure-as-code and automation tools (Terraform, Ansible, Python)
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes)
  • Demonstrated success in leading incident response and driving systemic improvements
  • Experience with capacity planning, performance tuning, and cost optimization
  • Excellent communication and stakeholder management skills, including executive engagement
  • Experience with ITIL/ITSM processes and integration with platforms like ServiceNow
  • Familiarity with security and compliance in regulated industries (financial services)
  • Background in performance engineering and infrastructure analytics
  • Experience developing dashboards and metrics for operational health and reliability

Bank of America Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Bank of America and has not been reviewed or approved by Bank of America.

  • Healthcare Strength Health coverage is described as comprehensive, with medical, dental, vision, virtual care via Teladoc, wellness programs, and specialized support for cancer and menopause. Wellness credits and an always‑on EAP with in‑person sessions add to the depth of care.
  • Parental & Family Support New parents can access up to 26 weeks of leave, including 16 weeks fully paid for eligible teammates, alongside back‑up child and adult care. Family‑building resources and reimbursements (e.g., fertility, adoption, surrogacy) and a dedicated Life Event Services team extend support across life stages.
  • Equity Value & Accessibility Broad‑based equity through the Sharing Success program, including $1B in stock to nearly all non‑executive employees in January 2026, is intended to foster an ownership mindset. Stock awards (including RSUs) are a recurring component that aligns employees’ interests with shareholders.

Bank of America Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Charlotte, NC
208,000 Employees
Year Founded: 1784

What We Do

We make financial lives better for our clients and our communities through the power of every connection. Our employees are at the heart of this purpose, and are key to driving responsible growth. Every day, across the globe, our employees bring a commitment to our purpose and to driving responsible growth by living our values: deliver together, act responsibly, realize the power of our people and trust the team. A key aspect of driving responsible growth is doing so in a sustainable manner, a critical pillar of which is being a great place to work for our teammates.

Gallery

Gallery

Similar Jobs

Hybrid
Houston, TX, USA
289097 Employees
Hybrid
Plano, TX, USA
289097 Employees
Hybrid
Plano, TX, USA
289097 Employees
Hybrid
Plano, TX, USA
289097 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account