Join us at the center of a rapidly growing technology field, where your work helps modernize complex, mission-critical systems. We’ll value your ideas, support your growth, and empower you to make reliability improvements that matter. Guidelines.docx Raw Posting.docx
Job summary
As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, you solve broad business problems with simple, straightforward solutions while improving the availability, reliability, and scalability of your application or platform. You use code and cloud infrastructure to configure, maintain, monitor, and optimize applications and their associated infrastructure, and you contribute meaningfully by sharing end-to-end operational knowledge across the team. We work collaboratively, communicate clearly during incidents, and focus on iterative improvements that reduce toil and improve outcomes.
Job responsibilities
- Design appropriate-level reliability designs, guide and assist others, and build consensus with peers while supporting adoption of site reliability engineering best practices within your team.
- Collaborate with software engineers and partner teams to design, develop, test, and implement deployment and reliability approaches using automated continuous integration and continuous delivery (CI/CD) pipelines.
- Implement infrastructure, configuration, and network as code for the applications and platforms in your remit, and iteratively improve solutions by decomposing problems into smaller, actionable changes.
- Operate and optimize applications and their associated infrastructure by configuring, maintaining, and monitoring services to meet availability, reliability, and scalability expectations.
- Resolve complex problems with technical experts, key stakeholders, and team members by using service level indicators (SLIs) and service level objectives (SLOs) to proactively address issues before they impact customers.
- Use enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Identify patterns in operational signals that indicate reliability risk or recurring toil, prioritize reuse-first improvements tied to SLO outcomes, and recognize roadblocks while exploring new technologies where appropriate.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience (country-specific requirements apply: NAMR/APAC—India/LATAM/Hong Kong; EMEA/LATAM—Brazil; Singapore follows local country guidance).
- Proficiency in site reliability engineering culture and principles, including how to implement site reliability engineering within an application or platform.
- Proficiency in at least one programming language such as Python, Java/Spring Boot, and .NET, with experience developing, debugging, and maintaining code in a large corporate environment.
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support site reliability engineering workflows, including strong validation habits and awareness of data sensitivity.
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements.
- Experience with observability practices such as white-box and black-box monitoring, SLO alerting, and telemetry collection, with familiarity troubleshooting common networking technologies and issues.
- Experience with continuous integration and continuous delivery tooling, plus familiarity with containers and container orchestration.
Preferred qualifications, capabilities, and skills
- Experience in site reliability engineering / production support / DevOps / platform roles with real on-call exposure, including improving reliability through incident response, root-cause analysis (RCA)/postmortems, problem management, and measurable toil reduction (reduced pages, automated repetitive tasks, improved MTTR/MTBF); ability to lead parts of incident calls, write clear RCAs, mentor others, and drive operational standards (runbooks, alerts, SLO definitions).
- Strong observability experience aligned to your stack, including Dynatrace (APM, alerting, dashboards) and/or Grafana + Prometheus; log analysis in Splunk (queries, dashboards, troubleshooting); ability to correlate metrics, logs, and traces to isolate issues in microservices; plus advanced observability concepts such as OpenTelemetry/distributed tracing, alert-as-code, SLO tooling, synthetic monitoring, and capacity planning using telemetry.
- Hands-on Kubernetes operations/troubleshooting (deployments, services/ingress, configmaps/secrets, HPA, node/pod debugging) and solid AWS experience supporting containerized workloads (EKS preferred where applicable, plus IAM/VPC basics); experience with CI/CD and infrastructure as code (Terraform modules, state management, pipeline integration; Jenkins/GitLab CI; blue/green and canary deployments) and advanced platform tooling (Helm/Kustomize, service mesh such as Istio/Linkerd, policy such as OPA/Gatekeeper, secrets tools such as Vault, GitOps such as ArgoCD/Flux).
- Strong automation mindset and microservices reliability depth: Java microservices plus scripting (Python and/or shell) to automate operational tasks (self-healing, runbook automation, unit tests) while following secure coding practices; Linux/Unix fundamentals (process/memory/disk troubleshooting, networking basics, log analysis, performance triage); operational experience supporting Kafka and/or MQ (including IBM MQ), and databases (Oracle and/or MongoDB) with awareness of performance symptoms and connection pools; certificate management in distributed systems (TLS/SSL, rotation/renewal, keystores/truststores, outage prevention); understanding of microservice failure modes (retries/timeouts, circuit breakers, rate limiting, backpressure, dependency mapping);
Skills Required
- Formal training or certification in site reliability engineering concepts
- At least 3 years of applied site reliability engineering experience
- Knowledge of SRE culture and principles
- Proficiency in at least one programming language, such as Python, Java/Spring Boot, or .NET
- Experience developing, debugging, and maintaining code in a large corporate environment
- Working knowledge of enterprise-authorized AI capabilities for SRE workflows
- Ability to validate AI-assisted operational recommendations and follow data-sensitivity requirements
- Experience with white-box and black-box monitoring, SLO alerting, and telemetry collection
- Familiarity troubleshooting networking technologies and issues
- Experience with continuous integration and continuous delivery tooling
- Familiarity with containers and container orchestration
- Production support, DevOps, platform, or SRE experience with on-call exposure
- Experience with incident response, root-cause analysis, postmortems, problem management, and toil reduction
- Experience leading incident calls, writing RCAs, mentoring others, and developing operational standards
- Experience with Dynatrace and/or Grafana and Prometheus
- Experience with Splunk log analysis, dashboards, and troubleshooting
- Experience correlating metrics, logs, and traces in microservices
- Knowledge of OpenTelemetry, distributed tracing, alert-as-code, SLO tooling, synthetic monitoring, and telemetry-based capacity planning
- Hands-on Kubernetes operations and troubleshooting
- AWS experience supporting containerized workloads, preferably EKS, including IAM and VPC basics
- Experience with Terraform, state management, and infrastructure-as-code pipeline integration
- Experience with Jenkins or GitLab CI and blue-green or canary deployments
- Experience with Helm, Kustomize, service mesh, policy tools, secrets management, or GitOps
- Java microservices and Python or shell scripting for operational automation
- Secure coding experience
- Linux or Unix troubleshooting fundamentals
- Operational experience supporting Kafka, MQ, or IBM MQ
- Experience with Oracle and/or MongoDB performance symptoms and connection pools
- Certificate management experience involving TLS/SSL, rotation, renewal, keystores, and truststores
- Understanding of microservice failure modes including retries, timeouts, circuit breakers, rate limiting, backpressure, and dependency mapping
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery







