You belong to the top echelon of talent in your field. At one of the world's most iconic financial institutions, where infrastructure is of paramount importance, you can play a pivotal role in shaping the reliability and resilience of systems that matter.
As a Lead Site Reliability Engineer at JPMorganChase within the Corporate Sector – Infrastructure Platforms, you hold a leadership role on your team, demonstrating strong knowledge across multiple technical domains and advising others on the technical and business challenges they face. You will lead resiliency design reviews, break complex problems into digestible work for other engineers, act as a technical lead for medium to large-scale products, and provide mentorship that elevates the entire team.
Job responsibilities
- Model and champion site reliability culture and practices; document and share knowledge across your organization through internal forums and communities of practice
- Lead initiatives to improve the reliability and stability of applications and platforms using data-driven analytics to improve service levels, proactively identifying and resolving technology-related bottlenecks
- Drive collaboration with your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets
- Serve as the primary point of contact during major incidents, applying deep technical expertise to identify and resolve issues quickly and minimize business impact
- Apply enterprise-authorized AI capabilities to accelerate major-incident triage, troubleshooting, post-incident analysis, and capacity risk identification — validating outputs and handling operational data according to sensitivity and security requirements
- Demonstrate an AI-first mindset by building and championing agentic automation for operational workflows (e.g., triage, runbook execution, incident summarization, change validation) with appropriate controls and monitoring
- Lead reuse-first adoption of AI-assisted reliability workflows across the software development lifecycle and toolchain practices, ensuring traceability, auditability, resiliency, and security controls
- Collect and analyze monitoring and telemetry data across test and production environments; design and implement dashboards to ensure system reliability, performance, and security
- Escalate issues with detailed technical write-ups; partner with application and infrastructure teams to identify and remediate capacity risks and understand platform interdependencies
- Offer a high level of technical expertise within one or more technical domains and provide advice and mentorship to other engineers
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Demonstrated proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and site reliability best practices, with hands-on experience implementing these within a platform
- Fluency in at least one programming language (e.g., Python, Java/Spring Boot, .NET, Go, Shell Scripting), including hands-on experience applying AI-assisted automation and agentic patterns to engineering and operational workflows
- Proficiency in scripting and automation and infrastructure-as-code tools (e.g., Python, PowerShell, Ansible, Terraform) and experience with cloud technologies across public and private environments
- Demonstrated experience using enterprise-authorized AI capabilities to improve site reliability engineering workflows (e.g., incident investigation, knowledge capture, capacity analysis) with strong validation habits and awareness of data sensitivity
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
- Advanced knowledge and experience in observability, monitoring, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Datadog, Prometheus, CloudWatch, or Splunk — including designing and implementing effective production monitoring dashboards
- Proficiency with continuous integration and continuous delivery practices and tooling, as well as container and container orchestration technologies
- Experience troubleshooting common networking technologies and issues, with knowledge of infrastructure areas including operating systems (Linux/Windows), databases, and deployment practices
- Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, and a demonstrated drive to self-educate and evaluate new technologies
Preferred qualifications, capabilities, and skills
- Hands-on experience and certifications in AWS, Azure, GCP, or other cloud environments, with understanding of resiliency, scalability, observability, and monitoring
- Experience implementing CI/CD pipelines, conducting code reviews using GitHub, and building process automation with Python and scripting
- Experience utilizing Terraform or other infrastructure-as-code technologies for cloud resource management
- Proven ability to leverage GitHub Copilot or similar coding assistants to accelerate skill development and build AI agents or workflows that support infrastructure operations tasks (e.g., triage, runbook execution)
- Experience supporting complex, mission-critical applications involving multiple components across varying technical generations, with familiarity with modern front-end technologies
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans
Skills Required
- Formal training or certification on infrastructure engineering concepts and 3+ years applied experience
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity
- Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations
- Strong knowledge of one or more scripting languages
- Experience with multiple cloud technologies (operate in and migrate across public and private clouds)
- Proficiency in scripting/automation and infrastructure-as-code (Python, PowerShell, Ansible, Terraform) and fluency in at least one programming language (Python, Go, Shell, .NET), including hands-on AI-assisted automation and agentic patterns
- Proven hands-on ability to leverage GitHub Copilot and coding assistants to build AI agents/workflows for infrastructure operations tasks
- Knowledge of operating systems (Linux/Windows), networking terminology/protocols, databases, deployment practices, and automation
- Familiarity with web products like Apache, Tomcat, IIS, WebSphere/IHS
- Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, and toil reduction with demonstrated implementation experience
- Advanced experience in observability, monitoring, alerting, and telemetry (Grafana, Dynatrace, Datadog, Prometheus, CloudWatch, Splunk), including designing/implementing dashboards and production support
- Implementation of CI/CD pipelines, code reviews using GitHub, and process automation with Python and scripting
- Hands-on experience and certifications in AWS, Azure, GCP (cloud exposure and resiliency/observability understanding)
- Experience utilizing Terraform or other IaC technologies for cloud resource management
- Experience supporting complex, mission-critical applications and familiarity with modern front-end technologies
- Drive to expand infrastructure engineering knowledge across emerging technologies and domains
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery









