Join our dynamic Production Management team to strengthen technology resilience, improve service quality, and drive effective governance across the core services that support our business.
As a Problem Management Lead in Employee Platforms, you will serve as a subject matter expert for end-to-end problem management, with a primary focus on identifying root causes, reducing recurring incidents, improving service stability, and ensuring that remediation and closure decisions meet established risk, control, resiliency, and documentation standards.
Success in this role requires strong analytical thinking, sound judgment, attention to detail, and the ability to influence technology teams and business stakeholders through clear communication, effective challenge, and data-driven recommendations.
Job responsibilities
Serve as the subject matter expert for end-to-end problem management across application and infrastructure services, from problem identification and classification through root cause analysis, remediation, validation, and closure.
Leads team adoption of enterprise-authorized AI capabilities within the work environment to improve incident triage speed and consistency (e.g., synthesizing operational signals into prioritized actions), with human-in-the-loop validation and appropriate handling of sensitive data.
Execute policies and procedures that ensure operational stability and availability and lead incident, problem, and change management in support of full stack technology systems, applications, or infrastructure.
Analyze incident trends, recurring issues, service-impacting events, and operational data to identify systemic problems and prioritize remediation based on business impact, risk, and customer experience.
Facilitate structured root cause analyses and post-incident reviews, ensuring contributing factors, control gaps, and lessons learned are documented and translated into sustainable corrective and preventive actions.
Establish clear ownership, milestones, and success measures for problem records and remediation actions; monitor progress, escalate delays, and provide transparent reporting to technology and business stakeholders.
Partner with engineering, operations, service owners, and control functions to deliver sustainable remediation, improve service resilience, and reduce the recurrence and business impact of known issues.
Maintain problem records and known-error information, ensuring workarounds, technical findings, risk acceptances, and remediation plans remain current, accessible, and aligned with established standards.
Identify operational risks and control deficiencies through problem management activities and coordinate appropriate mitigation, governance, audit, and compliance actions.
Define and report problem management metrics, including recurrence, aging, remediation effectiveness, root cause categories, and reductions in incident volume or business impact.
Applies reuse-first, AI-assisted practices across incident/problem/change routines to identify recurring interruption patterns and validate remediation actions aligned to resiliency and security expectations.
Required qualifications, capabilities, and skills
5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services.
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows with strong validation habits and awareness of data sensitivity.
Ability to review and validate AI-assisted incident recommendations before action, escalating when uncertain and ensuring outcomes align to operational, security, and auditability expectations.
Experience managing applications or infrastructure in a large-scale technology environment both on premises and public cloud.
Proficient in observability and monitoring tools and techniques.
Experience executing on processes in scope of the Information Technology Infrastructure Library (ITIL) framework.
Demonstrated experience managing end-to-end problem records, root cause analyses, corrective actions, known errors, and recurrence-prevention activities.
Strong analytical skills with experience using incident trends and operational data to identify systemic problems and prioritize remediation based on impact and risk.
Ability to manage operational risk, control requirements, resiliency expectations, and auditability needs within a complex technology environment.
Proven stakeholder management and influencing skills, including the ability to challenge constructively and communicate effectively with senior business and technology stakeholders.
Preferred qualifications, capabilities, and skills
Experience supporting Problem Management, Production Management, Site Reliability Engineering, or Operational Excellence activities within a large enterprise environment.
Knowledge of structured root cause analysis methods, post-incident review practices, known-error management, and corrective and preventive action management.
Experience defining and reporting problem management metrics, including recurrence, aging, remediation effectiveness, root cause categories, and service impact reduction.
Familiarity with enterprise-authorized AI, automation, analytics, or AIOps capabilities used to support operational analysis and workflow efficiency.
Relevant industry certification or formal training in ITIL, Service Management, Site Reliability Engineering, cloud technology, or operational risk management.
About Us
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans
Skills Required
- 5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services.
- Experience using enterprise-authorized AI capabilities in production operations workflows, with strong validation habits and data-sensitivity awareness.
- Ability to review and validate AI-assisted incident recommendations before action and escalate uncertainty.
- Experience managing applications or infrastructure in large-scale on-premises and public-cloud environments.
- Proficiency with observability and monitoring tools and techniques.
- Experience executing processes within the ITIL framework.
- Experience managing end-to-end problem records, root cause analyses, corrective actions, known errors, and recurrence prevention.
- Strong analytical skills using incident trends and operational data to identify systemic problems and prioritize remediation.
- Ability to manage operational risk, control requirements, resiliency expectations, and auditability needs.
- Stakeholder management and influencing skills, including constructive challenge and communication with senior stakeholders.
- Experience supporting Problem Management, Production Management, Site Reliability Engineering, or Operational Excellence in a large enterprise.
- Knowledge of structured root cause analysis, post-incident reviews, known-error management, and corrective and preventive action management.
- Experience defining and reporting problem management metrics.
- Familiarity with enterprise-authorized AI, automation, analytics, or AIOps capabilities.
- Relevant certification or formal training in ITIL, Service Management, Site Reliability Engineering, cloud technology, or operational risk management.
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery






