Vice President - Senior Manager of Site Reliability Engineering

Posted Yesterday
Be an Early Applicant
Jersey City, NJ, USA
Hybrid
Senior level
Financial Services
We’re one of the world’s biggest technology-driven companies
The Role
Leads a team of 8–10 site reliability and platform engineers supporting large-scale data and AI/ML platforms. Owns reliability requirements, availability targets, SLI/SLO frameworks, observability, incident management, infrastructure stability, CI/CD, Terraform, and automation. Oversees Databricks and Spark environments, promotes secure enterprise AI usage, manages stakeholders and compliance, and contributes to staffing, budgets, technical strategy, and engineering culture.
Summary Generated by Built In

Elevate your engineering leadership to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. In this high-impact role, you will guide and shape the future of large-scale data platform reliability — bringing your expertise in site reliability engineering, platform engineering, and AI/ML infrastructure to mentor and lead a team of 8–10 engineers.

As a Senior Manager of Site Reliability Engineering at JPMorganChase within the Chief Data and Analytics Office AI/ML and Data Platforms team, you are the non-functional requirement owner and champion for the applications in your remit. You will define availability targets, embed reliability principles into product design and testing, and ensure service level indicators and objectives are implemented in production to support secure, scalable, and high-performing analytics and AI/ML workloads. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.

Job responsibilities

  • Lead, mentor, and develop a team of 8–10 site reliability and platform engineers, fostering a culture of ownership, blameless post-mortems, and continuous improvement through tailored feedback and growth plans
  • Own and champion non-functional requirements, availability targets, service level indicators, and service level objectives for services supporting large-scale data platforms and AI/ML workloads, ensuring alignment with stakeholders and production readiness standards
  • Drive the design, implementation, and evolution of observability and reliability frameworks across distributed systems and data platform environments, leveraging tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk
  • Oversee the architecture and operational stability of data platform infrastructure, including Databricks, Spark-based data pipelines, and big data ecosystem tools, ensuring scalability, security, and high performance
  • Lead reuse-first adoption of enterprise-authorized AI capabilities within site reliability engineering workflows, establishing team standards for traceability, auditability, and alignment to resiliency and security expectations, with human-in-the-loop validation
  • Drive a culture of continual improvement by encouraging real-time feedback loops, conducting regular team debriefs, and applying objective, data-driven post-mortem strategies that enable teams to learn from both successes and failures
  • Manage stakeholders and ensure teams deliver projects aligned with compliance standards, risk and security requirements, service level agreements, and business objectives
  • Contribute to staffing, budget, and resource planning decisions, including hiring, developing, and recognizing engineering talent across the team
  • Champion site reliability engineering culture and principles across the organization, ensuring teams document and share knowledge and innovations via internal communities of practice, guilds, and engineering forums
  • Establish and govern CI/CD pipelines, infrastructure as code practices such as Terraform, and automation frameworks to accelerate delivery and reduce operational toil across the platform

Required qualifications, capabilities, and skills

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience. In addition, 2+ years of experience leading technologists to manage and solve complex technical items within your domain of expertise
  • Demonstrated experience managing and growing site reliability or platform engineering teams, with direct responsibility for a team of 8 or more engineers
  • Advanced proficiency in site reliability culture and principles, with a proven track record of implementing SLI/SLO/SLA frameworks, error budgets, incident management, and production readiness practices across large-scale distributed systems and data platforms
  • Hands-on experience with observability tooling and monitoring strategies, including white and black box monitoring, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk
  • Experience leading platform engineering efforts on large-scale data platforms and data lake ecosystems, including distributed compute frameworks such as Spark and managed platforms such as Databricks
  • Proficiency in Python or similar programming languages for automation, platform development, and operational tooling
  • Experience with containerization and orchestration technologies including Docker and Kubernetes, and infrastructure as code tools such as Terraform
  • Demonstrated experience leading teams in the safe and effective use of enterprise-authorized AI capabilities within reliability engineering workflows, including validation practices, escalation paths, and awareness of data sensitivity
  • Proficient with CI/CD practices and related tooling, container orchestration, and troubleshooting common networking technologies and issues
  • Strong communication and stakeholder management skills, with the ability to translate complex technical concepts into business-aligned outcomes and influence peers and executive partners

Preferred qualifications, capabilities, and skills

  • Experience with AWS platforms and cloud-native infrastructure services supporting AI/ML and analytics workloads at scale
  • Familiarity with big data ecosystem tools such as Spark, Glue, or MapReduce, and experience supporting data engineering teams in production environments
  • Experience building and managing CI/CD pipelines and automation frameworks that support platform reliability and engineering velocity
  • Background in AI/ML platform engineering, including infrastructure support for model training, serving, and monitoring pipelines
  • Demonstrated contributions to engineering communities through internal forums, communities of practice, or external conferences
About Us
JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. 

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans

About the TeamOur professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.

Skills Required

  • Formal training or certification in site reliability engineering concepts
  • 5+ years of applied site reliability engineering experience
  • 2+ years leading technologists and solving complex technical problems
  • Experience managing and growing site reliability or platform engineering teams
  • Direct responsibility for a team of 8 or more engineers
  • Experience implementing SLI, SLO, SLA, error budget, incident management, and production readiness practices
  • Hands-on experience with observability, monitoring, alerting, and telemetry tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk
  • Experience leading platform engineering for large-scale data platforms and data lake ecosystems
  • Experience with distributed compute frameworks such as Spark and managed platforms such as Databricks
  • Proficiency in Python or similar programming languages
  • Experience with Docker and Kubernetes
  • Experience with infrastructure as code tools such as Terraform
  • Experience leading safe and effective use of enterprise-authorized AI capabilities in reliability engineering workflows
  • Proficiency with CI/CD practices, container orchestration, and networking troubleshooting
  • Strong communication and stakeholder management skills
  • Experience with AWS platforms and cloud-native infrastructure services
  • Familiarity with Spark, Glue, MapReduce, or similar big data ecosystem tools
  • Experience supporting data engineering teams in production environments
  • Experience building and managing CI/CD pipelines and automation frameworks
  • Background in AI/ML platform engineering, including model training, serving, and monitoring infrastructure
  • Contributions to engineering communities, internal forums, communities of practice, or external conferences

JPMorganChase Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.

  • Healthcare Strength — Health coverage is considered comprehensive, including medical, dental, and vision, alongside wellness and mental health resources. Some locations add onsite health centers and related wellbeing support.
  • Retirement Support — Retirement offerings include a 401(k)-type savings plan and related financial benefits, with options such as employee stock purchase participation. Financial planning resources are also highlighted to support long-term savings.
  • Parental & Family Support — Paid parental leave of 16 weeks for birth or adoption is available for all parents. Child care and back-up child care resources further reinforce family support.

JPMorganChase Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
289,097 Employees
Year Founded: 1799

What We Do

JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.

Why Work With Us

Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.

Gallery

Gallery

Similar Jobs

Eve Logo Eve

Revenue Enablement Manager, Enterprise & Strategic Sales

Legal Tech • Software • Generative AI
Easy Apply
Remote or Hybrid
United States
180 Employees
Easy Apply
Remote or Hybrid
United States
180 Employees

DraftKings Logo DraftKings

Senior Manager, Lottery Enablement

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Remote or Hybrid
United States
6400 Employees
118K-148K Annually

DraftKings Logo DraftKings

Software Architect

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Remote or Hybrid
United States
6400 Employees
185K-232K Annually

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account