Senior Director of Site Reliability Engineering - Executive Director

Reposted 4 Hours Ago
Be an Early Applicant
London, Greater London, England, GBR
Hybrid
Senior level
Financial Services
We’re one of the world’s biggest technology-driven companies
The Role
Leads the global network estate SRE organization, defining reliability strategy, embedded team structures, SLOs, observability standards, incident practices, and toil-reduction automation. Builds and develops managers and senior engineers while driving AI-native engineering with validation, security, and operational-risk controls. Partners across network engineering, platform, data, and security teams to deliver reliable enterprise infrastructure in a regulated environment.
Summary Generated by Built In
Join an iconic company and take your career to new heights by leading talented teams in transformative projects. Together, let's push boundaries and achieve unparalleled success. 
As a Senior Director of Site Reliability Engineering at JPMorgan Chase in Infrastructure Platforms, you own and lead the SRE organization for our global network estate. You build and run teams of code-first SREs embedded within the network engineering teams, making reliability an engineered property of every platform rather than an operational afterthought. You are accountable for the reliability of enterprise-scale infrastructure: the SLIs/SLOs it is held to, the observability that proves it, the incident leadership that protects it, and the automation that engineers away toil. You lead an AI-native organization, driving fluent use of AI across the software development and operations lifecycle while holding the line on correctness, security, and risk.

Job Responsibilities

  • Owns the reliability of the network estate end-to-end: define the SRE operating model, the embedded-team structure, and the reliability outcomes the organization is accountable for
  • Builds, leads, and grows a multi-team SRE organization: hire and develop SRE managers and senior individual contributors, and set a high, consistent code-first engineering bar across teams
  • Runs the embedded model: place SRE teams inside the network engineering teams so reliability is engineered in at the source, while maintaining a coherent central discipline, shared standards, and career path
  • Establishes and govern SLIs, SLOs, and error budgets across platforms; drive SLO-based alerting, telemetry standards, and actionable observability as organizational defaults
  • Sets the strategy for toil reduction and self-healing infrastructure: treat repeated manual work as a defect to be engineered out, and hold teams to measurable reduction
  • Acts as the bridge between the network engineering teams and the automation-platform and data-engineering teams: translate operational reliability needs into platform and data requirements, and ensure the resulting tooling and data actually meet the business need and are adopted in production
  • Owns the major-incident and post-incident practice: strong incident leadership, blameless post-incident reviews, and durable engineering fixes that actually land
  • Drives an AI-native way of working across the org (AI-assisted development, code review, test generation, incident and root-cause analysis) with clear validation standards, so speed never compromises correctness, security, or reliability
  • Partners with network engineering, platform, and security leadership to align reliability strategy, roadmaps, and investment, and make the case for reliability work in business terms
  • Applies security and operational-risk judgment across the engineering lifecycle and ensure the organization operates within a regulated-enterprise control environment
  • Leads firmwide reuse-first adoption of enterprise-authorized AI capabilities within the work environment to accelerate reliability planning, operational learning, and delivery execution, with human-in-the-loop validation and appropriate handling of sensitive data

Required Qualifications, Capabilities, and Skills

  • Extensive experience leading SRE, production engineering, or reliability-focused software organizations at scale, including leading other managers (a leader of leaders)
  • A code-first foundation: credible software engineering background and the judgment to hold a code-first SRE bar, not an operations-only one
  • Proven track record running production systems at scale, including SLI/SLO/error-budget practice, incident leadership, and measurable toil reduction
  • Demonstrated ability to build and scale teams: hiring, developing SRE leaders and senior engineers, and establishing engineering culture and career paths
  • Demonstrated experience leading safe adoption of enterprise-authorized AI capabilities within the work environment at firm scale, including validation practices, data sensitivity considerations, and measurable reliability outcomes
  • Deep observability and reliability expertise: white-box/black-box monitoring, SLO-based alerting, and telemetry, and the ability to set these as organizational standards
  • Strong systems thinking across interfaces, contracts, failure modes, and interactions at enterprise scale
  • Fluency and conviction in AI-native engineering: directing AI to do real engineering and operations work, with sound judgment on where it applies and where deep human expertise is required
  • Security-first mindset and sound operational-risk judgment from design through production
  • Able to influence across a large, matrixed organization and lead calmly under pressure during high-severity events
  • Outcome orientation: focused on reliability, impact, and cost, not activity or span of control

Preferred Qualifications, Capabilities, and Skills

  • Networking depth (routing, switching, security, packet/flow analysis) or experience leading reliability for network or network-adjacent platforms: a strong plus, not a requirement
  • Experience running an embedded SRE model, driving reliability into engineering teams from within rather than from a central operations silo
  • Experience across multiple infrastructure domains
  • Demonstrated use of AI to redesign engineering and operational workflows for measurable impact, and to build organizational AI fluency
  • Experience supporting mission-critical systems in a production environment
  • Prior experience in regulated or large-scale enterprise environments

 

 

About UsJ.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives.
  
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
About the TeamOur professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.

Skills Required

  • Extensive experience leading SRE, production engineering, or reliability-focused software organizations at scale, including leading managers
  • Credible software engineering background and ability to maintain a code-first SRE standard
  • Experience operating production systems at scale using SLI, SLO, error-budget, incident leadership, and toil-reduction practices
  • Experience hiring and developing SRE leaders and senior engineers and establishing engineering culture and career paths
  • Experience leading safe adoption of enterprise-authorized AI capabilities, including validation practices and data-sensitivity considerations
  • Deep expertise in observability, white-box and black-box monitoring, SLO-based alerting, and telemetry
  • Strong systems-thinking skills across interfaces, contracts, failure modes, and enterprise-scale interactions
  • Experience directing AI-assisted engineering and operations work with sound judgment about human expertise requirements
  • Security-first mindset and operational-risk judgment throughout the engineering lifecycle
  • Ability to influence across a large, matrixed organization and lead calmly during high-severity incidents
  • Outcome-oriented approach focused on reliability, impact, and cost
  • Networking expertise in routing, switching, security, packet or flow analysis, or reliability leadership for network platforms
  • Experience operating an embedded SRE model within engineering teams
  • Experience across multiple infrastructure domains
  • Experience using AI to redesign engineering and operational workflows and build organizational AI fluency
  • Experience supporting mission-critical production systems
  • Experience in regulated or large-scale enterprise environments

JPMorganChase Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.

  • Healthcare Strength Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
  • Parental & Family Support Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
  • Retirement Support Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.

JPMorganChase Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
289,097 Employees
Year Founded: 1799

What We Do

JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.

Why Work With Us

Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.

Gallery

Gallery

Similar Jobs

Datadog Logo Datadog

Principal Partner Manager - Channels (UKI Security)

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
United Kingdom
6500 Employees

Flatiron Health Logo Flatiron Health

Reliability Engineer

Healthtech • Software • Biotech • Pharmaceutical
Hybrid
London, Greater London, England, GBR
2500 Employees

MongoDB Logo MongoDB

Senior Solutions Architect

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
London, Greater London, England, GBR
5550 Employees

CrowdStrike Logo CrowdStrike

Threat Analyst III (Remote, IRE)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
3 Locations
11000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account