As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.
Job responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
- Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
- Drive incident and problem management by running structured investigations, producing high-quality RCAs, and ensuring corrective/preventive actions are delivered.
- Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
- Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
- Engineer and support software-defined networking capabilities, including SD-WAN, SDA, and broader SND.
- Engineer and support routing and switching across enterprise networks.
- Engineer and support security and L4–L7 network components, including firewalls, load balancers, and proxies.
- Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
- Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
- Strong hands-on networking skills across enterprise routing/switching and security/L4–L7 components.
- Strong automation capability using Python, Shell, and Ansible in production operations.
- SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
- Ability to work independently, prioritize effectively, and deliver with minimal oversight.
- Experience supporting software-defined networking environments (e.g., SD-WAN, SDA, and related tooling).
- Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
- Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts.
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
- Demonstrate experience with Cisco ACI / fabrics.
- Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
- Work effectively in a financial institution or other regulated environment.
Skills Required
- Formal training or certification in site reliability engineering concepts and at least 3 years of applied experience
- Experience with incident response and problem management, including end-to-end RCA ownership through closure
- Strong hands-on enterprise networking skills across routing, switching, and security/L4-L7 components
- Production automation experience using Python, Shell, and Ansible
- Practical understanding of SRE concepts, non-functional requirements, and risk analysis methods such as FMEA
- Ability to work independently, prioritize effectively, and deliver with minimal oversight
- Experience supporting software-defined networking environments, including SD-WAN, SDA, and related tooling
- Ability to build and operationalize monitoring and observability through dashboards, alerting, and service health metrics
- Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts
- Working knowledge of enterprise-authorized AI capabilities for SRE workflows, including validation and data-sensitivity awareness
- Ability to validate AI-assisted operational recommendations before applying changes and escalate uncertainty appropriately
- Experience with Cisco ACI or fabrics
- CCNP, CCNA, or other relevant vendor certification
- Experience working in a financial institution or other regulated environment
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery






