Site Reliability Engineer III

Posted 16 Hours Ago
Be an Early Applicant
Hyderabad, Telangana, IND
Hybrid
Mid level
Financial Services
We’re one of the world’s biggest technology-driven companies
The Role
Develop and support scalable, resilient AI/ML data platforms; manage incidents, root cause analysis, production changes, observability, and disaster recovery. Build automation to reduce toil, apply authorized AI tools to improve incident response, and maintain SLO-driven reliability. Collaborate globally, mentor team members, manage operational risks, and contribute to strategic initiatives.
Summary Generated by Built In

Join a dynamic team where your expertise in site reliability engineering will shape the future of AI/ML data platforms. Unlock opportunities for growth and impact as you help build resilient, market-leading solutions.


As a Site Reliability Engineer III at JPMorgan Chase within the AI/ML Data Platforms team, you will play a pivotal role in developing scalable and resilient data solutions. You will engage in root cause analysis, production changes, and strategic initiatives that drive operational excellence. Your experience will help mentor team members and foster collaboration across global teams. Together, we create innovative solutions that support the firm’s mission and community.

 

Job responsibilities

  • Build and support scalable, resilient AI/ML data solutions
  • Coordinate incident management coverage for effective application issue resolution
  • Collaborate with cross-functional teams to perform root cause analysis and implement production changes
  • Develop and support AI/ML solutions for troubleshooting and incident resolution
  • Mentor and guide team members to drive strategic change
  • Manage budgetary considerations and staffing challenges
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Partner with colleagues across global teams to deliver impactful results

 

Required qualifications, capabilities and skills

  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Proficient in site reliability culture and principles, with familiarity in implementing site reliability within an application or platform
  • Proficiency in running production incident calls and managing incident resolution
  • Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
  • Strong understanding of SLI/SLO/SLA, Error Budgets and Proficiency in Python or PySpark for AI/ML modeling
  • Must be able to reduce toil by building new tools to automate repeated tasks
  • Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery
  • Awareness of risk controls and compliance with departmental and company-wide standards
  • Ability to work collaboratively in teams and build meaningful relationships to achieve common goals

 

Preferred qualifications, capabilities and skills

  • 4+ years in an SRE or production support role with AWS Cloud, Databricks, Snowflake or similar technologies
  • AWS and Databricks certifications
 

Skills Required

  • Formal training or certification in site reliability engineering concepts
  • 3+ years of applied site reliability engineering experience
  • Proficiency in site reliability culture and principles
  • Experience implementing site reliability within an application or platform
  • Experience running production incident calls and managing incident resolution
  • Experience with observability, including white-box and black-box monitoring, SLO alerting, and telemetry collection
  • Experience using enterprise-authorized AI capabilities in SRE workflows with validation and data-sensitivity awareness
  • Ability to validate AI-assisted operational recommendations and follow data-sensitivity requirements
  • Understanding of SLI, SLO, SLA, and error budgets
  • Proficiency in Python or PySpark for AI/ML modeling
  • Ability to reduce toil by building automation tools
  • Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery
  • Awareness of risk controls and departmental and company-wide compliance standards
  • Ability to collaborate effectively and build meaningful professional relationships
  • 4+ years in an SRE or production support role
  • Experience with AWS Cloud, Databricks, Snowflake, or similar technologies
  • AWS certification
  • Databricks certification

JPMorganChase Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.

  • Healthcare Strength Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
  • Parental & Family Support Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
  • Retirement Support Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.

JPMorganChase Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
289,097 Employees
Year Founded: 1799

What We Do

JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.

Why Work With Us

Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.

Gallery

Gallery

Similar Jobs

Optum Logo Optum

Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Software Engineering Lead

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Software Engineering Lead - Python Fullstack, FastAPI, GenAI, AI ML

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Principal Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account