As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Reliability Engineering & Problem Management team, you will serve as the technical authority for post-incident investigations and operational resilience. You will lead deep technical reviews, validate causal analysis, and ensure corrective actions address systemic causes. You will partner with cross-functional teams to drive measurable improvements in reliability, resilience, and operational excellence. You will help ensure incidents translate into lasting engineering improvements and reduction of repeat failures.
Job responsibilities
- Serve as the technical authority for Problem Management-led Root Cause Analysis reviews across major incidents and service-impacting events
- Lead deep-dive RCA challenge sessions, validating technical findings and causal chains through evidence-based analysis. Assess the quality and accuracy of root cause investigations, ensuring conclusions are technically sound and defensible. Challenge assumptions, unsupported conclusions, symptom-based findings, and ineffective corrective actions. Drive a culture of accountability focused on systemic learning and long-term reliability improvements
- Evaluate detection gaps, monitoring effectiveness, observability shortcomings, automation opportunities, resilience weaknesses, process breakdowns, and human factors. Validate that corrective actions address the true root cause and reduce recurrence likelihood and impact. Review corrective actions for closure and effectiveness, ensuring intended reliability and operational outcomes. Provide technical challenge and independent review of vendor, third-party, and internal investigation reports
- Use enterprise-authorized AI capabilities to accelerate reliability design and operational decisioning, validating outputs and handling operational data according to sensitivity and security requirements
- Lead reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices, ensuring traceability, auditability, resiliency, and security controls. Define and govern Service Level Objectives, Service Level Indicators, Reliability Metrics, and Error Budgets
- Evaluate system architecture against reliability design principles
- Recommend resilience patterns including graceful degradation, dependency isolation, rate limiting, circuit breakers, fault tolerance, capacity management, auto-remediation, and self-healing capabilities
- Lead reliability maturity assessments across platforms and services
- Lead enterprise-authorized AI capabilities for RCA generation, incident analysis, log analytics, pattern discovery, problem trend analysis, and corrective action recommendations
- Establish governance for explainability, auditability, data handling, and validation of AI-generated findings
- Drive AI-assisted reliability workflows across the incident lifecycle
Required qualifications, capabilities and skills
- Formal training or certification on security engineering concepts and 5+ years applied experience. Hands on experience in Infrastructure Engineering, Site Reliability Engineering, Production Engineering, Systems Engineering, or Software Engineering. Strong exposure to leading or supporting critical incident investigations
- Experience conducting deep technical RCAs for enterprise-scale environments, working in highly regulated and mission-critical environments
- Advanced expertise in Network Engineering, Cloud Infrastructure, Linux/Windows Platforms, Middleware Technologies, Database Technologies, Storage Platforms, Application Architecture, DevOps Toolchains, Distributed Systems, and Enterprise Monitoring Platforms
- Strong hands-on expertise in SLO/SLI Engineering, Distributed Tracing, Telemetry Design, Reliability Metrics, Error Budget Management, and AIOps Platforms
- Experience with tools such as Splunk, Dynatrace, Grafana, Datadog, Prometheus, AppDynamics, Elastic, and Open Telemetry
- Deep understanding of Root Cause Analysis methodologies, Five Whys, Fault Tree Analysis, Event Correlation, Human Factors Analysis, Systemic Cause Analysis, Problem Management Governance, and Major Incident Management
- Demonstrated experience using enterprise-authorized AI capabilities to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity
- Ability to set team practices for safe AI usage in operations while maintaining resiliency, security, and auditability outcomes
- Strong executive communication skills. Ability to challenge senior engineering stakeholders constructively. Proven ability to influence without direct authority
- Ability to translate technical findings into executive-ready narratives
Preferred qualifications, capabilities and skills
- Experience leading reliability engineering initiatives in large-scale, complex environments
- Expertise in AI-enabled incident analysis and reliability workflows
- Experience establishing governance for explainability and auditability of AI-generated findings
- Demonstrated ability to drive measurable improvements in reliability, resilience, and operational excellence
Skills Required
- Formal training or certification in security engineering concepts
- 5+ years of applied experience in security engineering concepts
- Hands-on experience in infrastructure engineering, site reliability engineering, production engineering, systems engineering, or software engineering
- Experience leading or supporting critical incident investigations
- Experience conducting deep technical root cause analyses in enterprise-scale environments
- Experience working in highly regulated and mission-critical environments
- Advanced expertise in network engineering, cloud infrastructure, Linux or Windows platforms, middleware, databases, storage, application architecture, DevOps toolchains, distributed systems, and enterprise monitoring
- Hands-on expertise in SLO/SLI engineering, distributed tracing, telemetry design, reliability metrics, error budgets, and AIOps platforms
- Experience with Splunk, Dynatrace, Grafana, Datadog, Prometheus, AppDynamics, Elastic, and OpenTelemetry
- Deep understanding of root cause analysis, Five Whys, fault tree analysis, event correlation, human factors analysis, systemic cause analysis, problem management governance, and major incident management
- Experience using enterprise-authorized AI capabilities to improve reliability engineering workflows
- Ability to establish safe AI usage practices while maintaining resilience, security, and auditability
- Strong executive communication skills and ability to influence without direct authority
- Ability to translate technical findings into executive-ready narratives
- Experience leading reliability engineering initiatives in large-scale, complex environments
- Expertise in AI-enabled incident analysis and reliability workflows
- Experience establishing governance for explainability and auditability of AI-generated findings
- Demonstrated ability to drive measurable improvements in reliability, resilience, and operational excellence
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery



.jpeg)





