Drive reliability at scale — join a team where your engineering expertise shapes the resilience of critical systems.
JPMorganChase is one of the world's leading financial services firms, and the technology that powers it demands the highest standards of reliability, performance, and scale. Here, you will work alongside talented engineers who are passionate about building systems that never sleep — and you will have the opportunity to grow your career while solving some of the most complex infrastructure challenges in the industry. We invest in our people, our platforms, and our future — and we want you to be part of it.
As a Lead Site Reliability Engineer at JPMorganChase, you will play a critical role in ensuring the availability, performance, and resilience of production systems that serve millions of customers and clients globally. You will partner with engineering and product teams to embed reliability practices into the software development lifecycle, driving a culture of operational excellence. Your work will directly impact the firm's ability to deliver seamless, uninterrupted services at enterprise scale.
Job responsibilities
- Lead the design and implementation of scalable, reliable, and observable infrastructure solutions that meet the firm's availability and performance standards
- Define and enforce service level objectives, error budgets, and reliability targets in partnership with engineering and product stakeholders
- Drive incident response, root cause analysis, and post-incident reviews to identify systemic improvements and reduce mean time to recovery
- Develop and maintain automation frameworks to eliminate toil, improve deployment pipelines, and accelerate delivery velocity
- Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process
- Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health
- Mentor and guide junior engineers, fostering a culture of continuous learning, operational discipline, and engineering excellence
- Evaluate and influence platform and tooling decisions to ensure alignment with reliability, scalability, and security requirements
- Leverage enterprise-authorized AI-assisted engineering practices to improve operational outcomes, including incident triage support, test strategy acceleration, and delivery workflow optimization, while ensuring consistent validation and secure handling of inputs and outputs
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and advanced applied experience
- Hands-on experience designing and operating large-scale distributed systems with a strong focus on availability, fault tolerance, and performance
- Proficiency in one or more programming or scripting languages (e.g., Python, Go, Java, Bash) for automation and tooling development
- Experience defining and managing service level indicators, service level objectives, and error budgets in production environments
- Strong background in observability tooling, including metrics, logging, and distributed tracing platforms
- Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements
- Experience with container orchestration and infrastructure-as-code practices (e.g., Kubernetes, Terraform, or equivalent)
- Ability to communicate complex technical concepts clearly to both technical and non-technical stakeholders
Preferred qualifications, capabilities, and skills
- Experience operating in a regulated financial services or similarly complex enterprise environment
- Familiarity with chaos engineering principles and tools used to proactively test system resilience
- Exposure to cloud-native architectures and multi-cloud or hybrid infrastructure environments
- Experience contributing to platform engineering or internal developer tooling initiatives
- Background in capacity planning, performance engineering, or cost optimization at scale
Skills Required
- Formal training or certification in site reliability engineering concepts
- Advanced applied experience in site reliability engineering
- Experience designing and operating large-scale distributed systems
- Strong focus on availability, fault tolerance, and performance
- Proficiency in one or more programming or scripting languages, such as Python, Go, Java, or Bash
- Experience defining and managing service level indicators, service level objectives, and error budgets in production environments
- Strong background in observability tooling, including metrics, logging, and distributed tracing platforms
- Experience leading incident response processes and conducting blameless postmortems
- Experience driving systemic reliability improvements
- Experience with container orchestration and infrastructure as code, such as Kubernetes or Terraform
- Ability to communicate complex technical concepts clearly to technical and non-technical stakeholders
- Experience in regulated financial services or a similarly complex enterprise environment
- Familiarity with chaos engineering principles and resilience-testing tools
- Exposure to cloud-native architectures and multi-cloud or hybrid infrastructure environments
- Experience contributing to platform engineering or internal developer tooling initiatives
- Background in capacity planning, performance engineering, or cost optimization at scale
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery







