Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Posted Yesterday
Be an Early Applicant
Chandler, AZ, USA
In-Office
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
The Role
Leads reliability and operational excellence for an enterprise Kubernetes container platform. Responsibilities include incident response, root cause analysis, L3 support, upgrades, observability, automation, capacity planning, performance tuning, security controls, compliance, resilience testing, and runbook development. The role partners with engineering, architecture, security, infrastructure, product, and operations teams while mentoring SRE resources and improving platform adoption, developer experience, and production reliability.
Summary Generated by Built In

Job Description:

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.
We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.
Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Position Summary:

The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:

Reliability & Operations

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Kubernetes & OpenShift Platform Engineering

  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption. 

Observability & Automation

  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.

Capacity & Performance Engineering

  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.

Security & Compliance

  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.

Required Qualifications:

Education / Experience

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Technical Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.

Desired Qualifications

  • BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications.
  • Experience running large-scale enterprise container platforms.
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
  • Experience implementing cloud-native security controls and platform governance.
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
  • Experience with vulnerability management and container security scanning solutions.
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicate effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.

Job Description:
This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices.

Responsibilities:

  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Participates regularly in architecture community of practice meetings and communication via other channels
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations

Skills:

  • Automation
  • Collaboration
  • Influence
  • Production Support
  • Result Orientation
  • Analytical Thinking
  • Application Development
  • Architecture
  • Solution Design
  • Stakeholder Management
  • Adaptability
  • DevOps Practices
  • Project Management
  • Risk Management
  • Solution Delivery Process

Shift:

1st shift (United States of America)

Hours Per Week: 

40

Skills Required

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience
  • 5+ years managing Kubernetes and/or OpenShift production environments
  • Experience operating large-scale mission-critical distributed systems
  • Experience supporting enterprise production environments with 24x7 operational responsibilities
  • Experience with Kubernetes, OpenShift, Rancher, and VKS container orchestration platforms
  • Linux administration and troubleshooting experience
  • Experience with Terraform, Ansible, GitOps, ArgoCD, and Helm
  • Experience with CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent
  • Experience with monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, or OpenTelemetry
  • Infrastructure-as-Code and automation framework experience
  • Knowledge of networking fundamentals, load balancing, ingress, DNS, and service mesh concepts
  • Experience with storage platforms, backup technologies, and disaster recovery solutions
  • Scripting experience in Python, Go, Bash, or similar languages
  • BS or MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience
  • OpenShift Administration or Kubernetes certification
  • Experience running large-scale enterprise container platforms
  • Experience with VMware, VCF, and OpenShift Virtualization
  • Experience implementing cloud-native security controls and platform governance
  • Experience with vulnerability management and container security scanning solutions

Bank of America Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Bank of America and has not been reviewed or approved by Bank of America.

  • Healthcare Strength Health coverage is described as comprehensive, with medical, dental, vision, virtual care via Teladoc, wellness programs, and specialized support for cancer and menopause. Wellness credits and an always‑on EAP with in‑person sessions add to the depth of care.
  • Parental & Family Support New parents can access up to 26 weeks of leave, including 16 weeks fully paid for eligible teammates, alongside back‑up child and adult care. Family‑building resources and reimbursements (e.g., fertility, adoption, surrogacy) and a dedicated Life Event Services team extend support across life stages.
  • Equity Value & Accessibility Broad‑based equity through the Sharing Success program, including $1B in stock to nearly all non‑executive employees in January 2026, is intended to foster an ownership mindset. Stock awards (including RSUs) are a recurring component that aligns employees’ interests with shareholders.

Bank of America Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Charlotte, NC
208,000 Employees
Year Founded: 1784

What We Do

We make financial lives better for our clients and our communities through the power of every connection. Our employees are at the heart of this purpose, and are key to driving responsible growth. Every day, across the globe, our employees bring a commitment to our purpose and to driving responsible growth by living our values: deliver together, act responsibly, realize the power of our people and trust the team. A key aspect of driving responsible growth is doing so in a sustainable manner, a critical pillar of which is being a great place to work for our teammates.

Gallery

Gallery

Similar Jobs

Byrd's House Detective DBA BHD Home and Mold Inspections Logo Byrd's House Detective DBA BHD Home and Mold Inspections

Virtual Assistant

Agency • Logistics • Real Estate • Virtual Reality • Consulting • Hospitality • Data Privacy
Remote or Hybrid
United States
100 Employees
35-45 Annually

Boeing Logo Boeing

Project Management Specialist (Mid-Level or Lead)

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Mesa, AZ, USA
170000 Employees
94K-160K Annually

Boeing Logo Boeing

Production Coordinator - Warehouse

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Mesa, AZ, USA
170000 Employees
44K-55K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
45K-85K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account