As a VP Infrastructure Engineer (Site Reliability Engineering), you will help build and mature a meaningful engineering discipline that combines software and systems to solve operational problems with scalable, automated, and resilient solutions. Your work will focus on optimizing existing systems, strengthening infrastructure reliability, and reducing manual work through automation across VMware Cloud Foundation (vCF) environments.
You will join the Global ESX Infrastructure Support team and leverage expertise in vSphere, NSX, and vCF platforms to drive measurable improvements in how we operate production applications and systems.
Key responsibilities- Lead reliability outcomes for critical production applications and infrastructure in large-scale, high-availability environments.
- Drive execution of service-level changes in VMware vCF environments, ensuring safe change practices, strong controls, and predictable outcomes.
- Troubleshoot complex multi-layer issues across compute, storage, virtualization, and network domains; coordinate incident response and drive restoration.
- Build automation to reduce toil and standardize operations (self-service where appropriate), improving operational efficiency and consistency.
- Perform analytics on incidents and usage patterns to identify trends, predict risk, and implement proactive remediation and prevention.
- Build and drive adoption of self-healing, resiliency, and “operability by design” patterns (runbooks, guardrails, monitoring/alerting, reliability reviews).
- Lead and participate in performance testing; identify bottlenecks, optimization opportunities, and capacity demands; influence roadmaps and investment decisions.
- Provide technical leadership across teams, translating reliability goals into actionable engineering work and measurable outcomes (availability, latency, error rates, MTTR).
- Mentor engineers, contribute to engineering standards, and help shape SRE practices across the organization.
- Applies reuse-first, AI-assisted practices across incident/problem/change routines to identify recurring interruption patterns and validate remediation actions aligned to resiliency and security expectations.
- Leads team adoption of enterprise-authorized AI capabilities within the work environment to improve incident triage speed and consistency (e.g., synthesizing operational signals into prioritized actions), with human-in-the-loop validation and appropriate handling of sensitive data
- 7+ years of experience in network engineering, software development, systems engineering, or related disciplines, with demonstrated progression in scope and ownership.
- Experience as a VMware ESXi system engineer in an enterprise-class environment, including vSAN.
- Proficiency troubleshooting components and executing service-level changes within VMware vCF environments.
- Experience operating in large-scale, dynamic, high-availability production environments.
- Software development/scripting experience in one or more general-purpose languages: Python, JavaScript, PowerShell.
- Strong computing fundamentals (systems, networking, troubleshooting, reliability mindset).
- Strong data center networking knowledge with protocols/technologies such as BGP, VRF, MPLS, IP, TCP, UDP.
- Demonstrated leadership in an agile team, with strong communication and the ability to partner effectively with both technical and non-technical audiences.
- Strong critical thinking and problem-solving skills; ability to work independently across multiple teams.
- Ability to review and validate AI-assisted incident recommendations before action, escalating when uncertain and ensuring outcomes align to operational, security, and auditability expectations.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows with strong validation habits and awareness of data sensitivity
- Advanced English skills
- VMware NSX-T (strong plus).
- Hyper-V (strong plus).
- Experience driving observability and reliability improvements (monitoring, alerting, incident/problem management, post-incident reviews).
- Familiarity with AI/GenAI-assisted engineering/operations (e.g., using AI to accelerate troubleshooting, automate remediation workflows, improve runbooks, or enhance incident analysis) is a plus.
Skills Required
- 7+ years experience in network engineering, software development, systems engineering, or related disciplines
- Experience as a VMware ESXi system engineer in an enterprise-class environment, including vSAN
- Proficiency troubleshooting and executing service-level changes within VMware Cloud Foundation (vCF) environments
- Experience operating in large-scale, dynamic, high-availability production environments
- Software development/scripting experience in Python, JavaScript, or PowerShell
- Strong computing fundamentals (systems, networking, troubleshooting, reliability mindset)
- Strong data center networking knowledge (BGP, VRF, MPLS, IP, TCP, UDP)
- Demonstrated leadership in an agile team and strong communication skills
- Ability to review and validate AI-assisted incident recommendations and ensure security/auditability
- Demonstrated experience using enterprise-authorized AI capabilities in production operations workflows
- Advanced English skills
- VMware NSX-T
- Hyper-V
- Experience driving observability and reliability improvements (monitoring, alerting, incident/problem management)
- Familiarity with AI/GenAI-assisted engineering/operations
JPMorganChase Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about JPMorganChase and has not been reviewed or approved by JPMorganChase.
-
Healthcare Strength — Medical, dental, vision, and mental-health coverage are broad, with wellness incentives, on-site or virtual care, and an EAP offering coaching and counseling. Plan materials emphasize accessible options, including multiple medical choices and tools to manage costs.
-
Parental & Family Support — Paid parental leave extends up to 16 weeks for all parents, supplemented by paid Critical Caregiver Leave. Family resources include backup childcare via Bright Horizons, lactation support and milk-shipping, family-building assistance, and even a free five-month SNOO rental for newborns.
-
Retirement Support — Retirement programs include a 401(k) with an annual company match and automatic pay credits for most employees, with a legacy pension available to earlier hires. An Employee Stock Purchase Plan at a 5% discount further supports long-term savings.
JPMorganChase Insights
What We Do
JPMorgan Chase & Co. (NYSE: JPM) is a leading global financial services firm with assets of $3.7 trillion and operations worldwide. The firm is a leader in investment banking, financial services for consumers and small businesses, commercial banking, financial transaction processing, and asset management. A component of the Dow Jones Industrial Average, JPMorgan Chase & Co. serves millions of consumers in the United States and many of the world’s most prominent corporate, institutional and government clients under its J.P. Morgan and Chase brands. Technology fuels every aspect of our company and is at the heart of everything we do. With over 50,000 technologists globally and an annual tech spend of $12 billion, we are dedicated to improving the design, analytics, development, coding, testing and application programming that goes into creating high quality software and new products. Learn more about technology at our firm, explore resources from our Distinguished Engineers, AI & ML researchers, and other experts; access the latest episode of our TechTrends podcast, and more at www.jpmorgan.com/technology. Information about JPMorgan Chase & Co. is available at www.jpmorganchase.com. ©2023 JPMorgan Chase & Co. All rights reserved. JPMorgan Chase is an Equal Opportunity Employer, including Disability/Veterans.
Why Work With Us
Our technologists work on a diverse range of solutions that include strategic technology initiatives, big data, mobile, electronic payments, machine learning, cybersecurity, enterprise cloud development, and other state-of-the-art technologies.
Gallery






