DESCRIPTION:
Duties: Define and enforce measurable reliability targets for critical application environments, ensuring operational metrics are consistently tracked and achieved. Architect, implement, and refine advanced monitoring and alerting solutions to proactively surface application health and performance issues across complex technology stacks. Collaborate with cross-functional teams to design and support robust, highly available application platforms capable of meeting stringent business and technical demands. Drive adoption of reliability and resilience best practices providing technical leadership to elevate operational standards across support and engineering groups. Develop, automate, and maintain failover and recovery workflows for application services ensuring uninterrupted operations across diverse infrastructure and cloud regions. Create, update, and operationalize incident response documentation and automated remediation mechanisms including rollback and circuit breaker strategies for application failures. Lead critical incident management for application platforms ensuring rapid restoration, comprehensive root cause analysis, and implementation of long-term improvements. Optimize resource allocation and infrastructure costs for large-scale application support, balancing efficiency with reliability, and performance requirements. Facilitate seamless integration and deployment of application changes bridging development and support to ensure smooth operational transitions. Oversee continuous validation processes including pre- and post-deployment monitoring to detect and remediate application drift and performance regressions. Maintain day-to-day operational stability and high availability for application systems, leveraging deep technical expertise in support and troubleshooting. Monitor production environments using advanced diagnostic and observability tools rapidly identifying and resolving anomalies. Escalate and communicate complex technical issues, delivering actionable insights and solutions to both technical and business stakeholders.
QUALIFICATIONS:
Minimum education and experience required: Bachelor's degree in Information Systems Engineering, Computer Engineering, or related field of study plus 5 years of experience in the job offered or as Site Reliability Engineer, Data Engineer, Data Analyst, MSSQL Server Developer, Support Engineer, Software Developer, or related occupation.
Skills Required: This position requires five (5) years of experience with the following: Utilizing Sev1/Sev2 on-call including triaging, mitigating, coordinating restoration, RCA, and blameless postmortems; MTTR recurrence reduction; Observability including metrics, logs, traces; low-noise alerts, fast detection; maintaining runbooks and escalations; Automation including Python and Bash; utilizing CI/CD with blue and green or canary, feature flags, pre-deploy validation, and reliable rollback; using IaC including Terraform for public cloud; using Modules, remote state, drift detection, LUT/policy-as-code, and targeted applies; using Kubernetes for deployments, autoscaling, health probes, progressive rollouts and rollbacks, quotas, and cluster troubleshooting; utilizing AWS production operations including VPC networking, IAM policy design, encryption/ KMS, load balancing, multi-AZ resilience, backup and restore, and regional failover; Database reliability including PostgreSQL, MySQL, Oracle backups/PITR, replication and failover, online schema changes, SQL and index tuning under load; Linux and networking including processing memory/IO diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery; Security including least privilege, secrets management, patch, vulnerability management, immutable audit logging and framework alignment. This position requires three (3) years of experience with the following: Utilizing SLI/SLO and error budgets sustaining ≥99.9% SLOs; Gate high-risk changes; Performance and capacity including load testing (JMeter, BlazeMeter), trace hot paths, tail-latency/throughput tuning, and peak demand modeling; Resilience and DR including timeouts, retries, backpressure, circuit breaking, graceful degradation; validated RTO/RPO; using Linux and networking to process memory/IO diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery; Releasing change management including versioned applications, DB migrations, automated quality gates Maxwell, controlled rollouts, and rapid clean NB rollback; Reliability outcomes including sustaining lower bast MT pipeline failure TR, fewer false alarms, higher S, logical improvements, safer deployments, hardened failure modes, and predictable peak scaling.
Job Location: 8181 Communications Pkwy, Plano, TX 75024.
Full-Time.
About UsWe offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans
Skills Required
- Bachelor's degree in Information Systems Engineering, Computer Engineering, or a related field
- Five years of experience in the job offered or as a Site Reliability Engineer, Data Engineer, Data Analyst, MSSQL Server Developer, Support Engineer, Software Developer, or related occupation
- Five years of experience with Sev1/Sev2 on-call incident response, triage, mitigation, restoration coordination, root cause analysis, and blameless postmortems
- Five years of experience reducing MTTR recurrence and implementing observability using metrics, logs, traces, low-noise alerts, runbooks, and escalation processes
- Five years of experience automating with Python and Bash
- Five years of experience with CI/CD, blue-green or canary deployments, feature flags, pre-deployment validation, and rollback
- Five years of experience with Terraform infrastructure as code for public cloud, including modules, remote state, drift detection, policy as code, and targeted applies
- Five years of experience using Kubernetes for deployments, autoscaling, health probes, progressive rollouts, quotas, and cluster troubleshooting
- Five years of experience with AWS production operations, including VPC networking, IAM policy design, KMS encryption, load balancing, multi-AZ resilience, backup and restore, and regional failover
- Five years of experience with PostgreSQL, MySQL, and Oracle database backups, point-in-time recovery, replication, failover, schema changes, SQL, and index tuning
- Five years of experience with Linux and networking diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery
- Five years of experience with least-privilege security, secrets management, patching, vulnerability management, immutable audit logging, and security framework alignment
- Three years of experience using SLI/SLOs and error budgets to sustain at least 99.9% SLOs and gate high-risk changes
- Three years of experience with load testing using JMeter or BlazeMeter, trace analysis, tail-latency and throughput tuning, and peak demand modeling
- Three years of experience implementing resilience and disaster recovery practices, including timeouts, retries, backpressure, circuit breaking, graceful degradation, and validated RTO/RPO
- Three years of experience with versioned application releases, database migrations, automated quality gates, controlled rollouts, and rapid rollback
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery








