Lead Site Reliability Engineer

Posted 6 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
Hybrid
Expert/Leader
Financial Services
We’re one of the world’s biggest technology-driven companies
The Role
Lead SRE transformation: define vision, operating model, and SLO/SLI frameworks; build a core SRE team; drive observability, incident response, resiliency patterns, automation, and AI-enabled reliability improvements across global production systems.
Summary Generated by Built In

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank – Markets – Sales, Research & Data Technology (SRDT), you hold a technical leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. You will will set the vision, strategy, and operating model for our SRE transformation—enabling our business-aligned support teams to deliver higher reliability, stronger resilience, and a measurably better end-user experience across the board. A key pillar of the role is driving our AI transformation: using AI responsibly and securely to reduce toil, strengthen observability, and move us from reactive operations to proactive and predictive reliability management. The successful candidate will inspire and motivate the team and will also be a hands-on engineer designing and delivering tangible solutions.


Job responsibilities

  • Defines the SRE vision, north-star outcomes, and multi-year roadmap for the Production Management team, aligned to both CIB and JPM Global Technology priorities.
  • Establishes the SRE operating model across global regions (ways of working, intake, prioritization, engagement with engineering teams and production support).
  • Partners with business-aligned Production Support leads to embed SRE practices consistently and act as a force multiplier – coaching them on reliability thinking, prioritization, and “engineering out” operational load.
  • Builds and develops a small, high-impact core SRE team (and/or virtual SRE community of practice) that scales reliability improvements across many application flows.
  • Defines and implements standards for: Service cataloging, SLO/SLI frameworks and error budgets, incident response maturity, blameless post-incident reviews, resiliency patterns, capacity, performance, and scalability engineering.
  • Drives service reviews with evidence-based reporting (availability, latency, incident trends, MTTR/MTTD, change failure rate, customer impact).
  • Champions AI adoption and deliver AI-enabled capabilities to reduce operational toil and improve speed/quality of response.
  • Sets direction for observability across logs/metrics/traces, including instrumentation standards, golden signals, and end-user journey monitoring. Improves alert quality and routing: reduce false positives, improve actionable alerts, and tighten feedback loops to engineering teams.
  • Builds strong partnerships with application development teams, platform/infrastructure partners, and governance functions. Communicates clearly and credibly at all levels—from engineers to senior technology and business stakeholders.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • 10+ years of experience in technology support, production/application support, DevOps, or infrastructure management.
  • Demonstrated experience leading SRE/reliability engineering or production engineering transformations in a complex enterprise environment.
  • Strong engineering background: ability to design, build, and deliver automation and reliability solutions.
  • Fluency & expertise in at least one programming language such as Python, Java Spring Boot or .Net
  • Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction.
  • Proficiency and experience with telemetry (logs/metrics/traces) collection using tools and standards such as Prometheus, OpenTelemetry, Datadog, Dynatrace, Splunk.
  • Experience delivering automation at scale (scripting, workflow automation, runbook automation, CI/CD-integrated guardrails).
  • Proven leadership skills: influencing without authority, coaching leaders, and building communities of practice.
  • Strong judgment around risk, security, and controls—especially when applying AI to production workflows.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.

Preferred qualifications, capabilities, and skills

  • Familiarity with ITIL support methodologies and concepts.
  • Solid understanding of networking concepts and troubleshooting.
  • Familiarity with Infrastructure as Code (IaC) tools and major cloud platforms (AWS, Azure, or GCP).

Skills Required

  • 10+ years of experience in technology support, production/application support, DevOps, or infrastructure management.
  • Demonstrated experience leading SRE/reliability engineering or production engineering transformations in a complex enterprise environment.
  • Strong engineering background: ability to design, build, and deliver automation and reliability solutions.
  • Fluency and expertise in at least one programming language such as Python, Java Spring Boot or .Net.
  • Deep practical knowledge of SLOs/SLIs, error budgets, incident management, postmortems, observability across metrics/logs/traces, resilience engineering, performance/capacity management, and change risk reduction.
  • Proficiency and experience with telemetry collection using tools and standards such as Prometheus, OpenTelemetry, Datadog, Dynatrace, Splunk.
  • Experience delivering automation at scale (scripting, workflow automation, runbook automation, CI/CD-integrated guardrails).
  • Proven leadership skills: influencing without authority, coaching leaders, and building communities of practice.
  • Strong judgment around risk, security, and controls, especially when applying AI to production workflows.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows with strong validation habits and data sensitivity awareness.
  • Ability to evaluate AI-assisted operational recommendations, define guardrails, and ensure outcomes meet resiliency and security expectations.

JPMorganChase Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.

  • Healthcare Strength Medical, dental, vision, and mental health coverage are comprehensive, with on-site clinics, preventive care, and specialized supports such as maternity nurse guidance and fertility treatments. Wellness activities can help offset copays and out-of-pocket costs, reinforcing the perceived strength of health benefits.
  • Retirement Support A 401(k) with dollar-for-dollar matching and additional automatic pay credits reflect strong employer-backed retirement savings. An employee stock purchase plan and related financial programs further bolster long-term financial support.
  • Leave & Time Off Breadth Paid time off, sick time, holidays, and generous parental leave are provided alongside family medical leave and adoption/fertility assistance. Additional programs like caregiver support and volunteer time off expand the breadth of time-away options.

JPMorganChase Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
289,097 Employees
Year Founded: 1799

What We Do

JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.

Why Work With Us

Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.

Gallery

Gallery

Similar Jobs

Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Optum Logo Optum

Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Fiserv Logo Fiserv

Site Reliability Engineer

eCommerce • Fintech • Information Technology • Payments • Financial Services
In-Office
Bengalurus, Bangalore, Karnataka, IND
41000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account