The Company
PayPal has been revolutionizing commerce globally for more than 25 years. Creating innovative experiences that make moving money, selling, and shopping simple, personalized, and secure, PayPal empowers consumers and businesses in approximately 200 markets to join and thrive in the global economy.
We operate a global, two-sided network at scale that connects hundreds of millions of merchants and consumers. We help merchants and consumers connect, transact, and complete payments, whether they are online or in person. PayPal is more than a connection to third-party payment networks. We provide proprietary payment solutions accepted by merchants that enable the completion of payments on our platform on behalf of our customers.
We offer our customers the flexibility to use their accounts to purchase and receive payments for goods and services, as well as the ability to transfer and withdraw funds. We enable consumers to exchange funds more safely with merchants using a variety of funding sources, which may include a bank account, a PayPal or Venmo account balance, PayPal and Venmo branded credit products, a credit card, a debit card, certain cryptocurrencies, or other stored value products such as gift cards, and eligible credit card rewards. Our PayPal, Venmo, and Xoom products also make it safer and simpler for friends and family to transfer funds to each other. We offer merchants an end-to-end payments solution that provides authorization and settlement capabilities, as well as instant access to funds and payouts. We also help merchants connect with their customers, process exchanges and returns, and manage risk. We enable consumers to engage in cross-border shopping and merchants to extend their global reach while reducing the complexity and friction involved in enabling cross-border trade.
Our beliefs are the foundation for how we conduct business every day. We live each day guided by our core values of Inclusion, Innovation, Collaboration, and Wellness. Together, our values ensure that we work together as one global team with our customers at the center of everything we do – and they push us to ensure we take care of ourselves, each other, and our communities.
Job Summary:
What do you need to know about the roleThis is an incident command role. You'll direct application and infrastructure teams during incidents making work assignments, prioritizing troubleshooting paths, and authorizing critical actions like rollbacks and regional failovers. You need the technical depth to rapidly read Infrastructure as Code, Kubernetes manifests, and CI/CD configurations to make informed decisions under pressure.
Meet Our Team
The Site Health Engineering (Command Center) team serves as the operational and technical authority during PayPal's most critical incidents. We are a team of experienced infrastructure and reliability professionals who blend deep technical expertise with sound judgment under pressure, directing cross-functional engineering efforts across PayPal's core platforms and family of brands, including Venmo, Xoom, Zettle, and Braintree.
Job Description:
Essential Responsibilities:
- Delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC) (design, implementation, testing, delivery and operations), based on definitions from more senior roles.
- Advises immediate management on project-level issues
- Guides junior engineers
- Operates with little day-to-day supervision, making technical decisions based on knowledge of internal conventions and industry best practices
- Applies knowledge of technical best practices in making decisions
Minimum Qualifications:
- 3+ years relevant experience and a Bachelor’s degree OR Any equivalent combination of education and experience.
Additional Responsibilities & Preferred Qualifications:
Your way to Impact
Rather than building or maintaining infrastructure day-to-day, our team is entrusted with a broader mandate: safeguarding the reliability, resiliency, and availability of some of the world's most heavily trafficked financial platforms. We hold final decision-making authority during high-severity incidents, partner closely with executive leadership on post-incident learnings, and drive the tooling and processes that continuously strengthen our incident response capabilities.
You'll also regularly interface with executive leadership during critical incidents and post-mortems, and drive implementation of tooling that advances the Command Center's capabilities.
Site Resiliency & Infrastructure Management
Proactively identify and address vulnerabilities in cloud (AWS, GCP, Azure) and on-premises infrastructure
Review Infrastructure as Code changes for reliability risks as part of change approval process
Identify architectural anti-patterns in Kubernetes deployments and cloud migrations
Conduct regular disaster recovery drills and readiness tests before major events (Thanksgiving, Cyber 5, peak shopping seasons)
Participate in situation room activities for new product rollouts
Drive site resilience projects to enhance system reliability and uptime
Proactively identify and address vulnerabilities in cloud (AWS, GCP, Azure) and on-premises infrastructure
Implement automated monitoring solutions to detect single points of failure
Lead new datacenter and CDN certification initiatives
Conduct regular disaster recovery drills and readiness tests before major events (Thanksgiving, Cyber 5, peak shopping seasons)
Participate in situation room activities for new product rollouts
Drive site resilience projects to enhance system reliability and uptime
Incident Management & Response
Act as incident commander with final decision authority -- directing engineering teams, authorizing rollbacks, and commanding regional failovers
Direct application and infrastructure teams during incidents by making work assignments and prioritizing troubleshooting paths
Rapidly assess incidents by reading Infrastructure as Code (Terraform, CloudFormation), Kubernetes manifests, and CI/CD configurations
Give final authorization for critical actions including production rollbacks, regional failovers, and emergency changes
Interface with executive leadership during critical incidents and post-mortems to provide technical guidance and impact assessments
Identify when incidents stem from teams deviating from established cloud-native patterns
Command cross-functional teams during high-severity incidents affecting PayPal core and brand platforms (Venmo, Xoom, Zettle, Braintree)
Lead blameless postmortem sessions and contribute to Root Cause Analysis (RCA) processes
Drive continuous improvement initiatives based on incident learningsServe as the primary technical escalation point during critical incidents
Accelerate incident response times through standardized playbooks and automated workflows
Coordinate cross-functional teams during high-severity incidents affecting PayPal core and brand platforms (Venmo, Xoom, Zettle, Braintree)
Lead blameless postmortem sessions and contribute to Root Cause Analysis (RCA) processes
Drive continuous improvement initiatives based on incident learnings
Manage multiple concurrent incidents during peak periods with efficiency and precision
Change Management & Risk Mitigation
Serve as final approver for emergency changes and provide expert guidance on all production changes
Act as advisor and technical authority during change approval processes, identifying potential reliability risks
Provide training and guidance to engineering teams on change management best practices
Maintain change audit documentation and compliance requirements
Review and approve changes to production systems, ensuring comprehensive risk assessment
Automate change validation and rollback procedures to minimize service disruptions
Streamline change management processes to reduce manual errors and bottlenecks
Provide training and guidance to engineering teams on change management best practices
Maintain change audit documentation and compliance requirements
Cloud Expertise & Technical Leadership
Leverage deep expertise in cloud platforms (AWS, GCP, Azure) to drive incident resolution
Support Braintree and Venmo cloud infrastructure operations
Guide teams toward solutions by providing architectural direction during incidents
Stay current with emerging cloud technologies and best practices
Mentor team members on cloud technologies and incident management techniques
Cloud Expertise & Technical Leadership
Implement automation, dashboards, and tooling to enhance the team's incident response capabilities
Build runbooks and playbooks for cloud-native incident scenarios
Develop internal tools and scripts to improve TDO operational efficiency
Drive projects that advance the Command Center's operational capabilities
In your day-to-day role you will be responsible for:
3days -4days alternating 12-hour shift pattern.
Take ownership of system performance monitoring, identify inefficiencies, and lead initiatives to improve the overall availability and reliability of digital platforms and applications.
Lead and manage the response to complex, high-priority incidents, ensuring prompt resolution and a thorough root cause analysis to prevent future occurrences.
Design and implement advanced automation frameworks to improve operational efficiency, streamline processes, and reduce human error.
Lead reliability-focused initiatives, ensuring systems are highly available, resilient, and scalable, and promote best practices across engineering teams.
Enhance the monitoring infrastructure by identifying key metrics, optimizing alerting, and improving system observability to ensure the reliability of large-scale systems.
Forecast resource requirements and lead capacity planning activities to ensure systems can scale effectively to meet growing user demand.
Ensure robust disaster recovery strategies are in place and conduct regular testing to ensure systems can recover quickly from failures.
Partner with engineering and product teams to identify opportunities for improving system architecture, focusing on scalability, reliability, and fault tolerance.
Provide mentorship and technical guidance to junior site reliability engineers, fostering skill development and knowledge sharing.
Drive continuous improvement across operational workflows, identifying areas for optimization, cost reduction, and performance enhancement.
What do you need to bring
Technical Skills
Significant hands-on experience with at least one major cloud provider (AWS or GCP required; multi-cloud experience preferred)
Strong proficiency with Infrastructure as Code tools (Terraform, CloudFormation, Pulumi, or equivalent) including ability to read, review, and troubleshoot IaC configurations during incidents
Significant hands-on experience with Kubernetes and CNCF ecosystem tools, including troubleshooting K8s deployments, manifests, and cluster issues
Ability to quickly read and review code across multiple languages (Python, Go, Bash) and configuration formats (YAML, HCL, JSON)essential for effective incident troubleshooting
Proven experience managing critical incidents in Infrastructure-as-Code driven environments, including troubleshooting IaC state issues, GitOps failures, and cloud-native deployment problems
Professional-level certification in at least one major cloud platform (AWS Solutions Architect Professional, Google Cloud Professional Cloud Architect, or equivalent)
Experience with monitoring and observability tools (Splunk, Datadog, Prometheus, Grafana)
Strong knowledge of networking, load balancing, CDN technologies, and DNS management
Proficiency in scripting for operational automation (Python, Bash, PowerShell)5+ years of experience in site reliability engineering, infrastructure operations, or similar technical operations roles
Strong expertise in cloud platforms (AWS, GCP, and/or Azure)
Proficiency in infrastructure automation tools (Terraform, Ansible, CloudFormation, etc.)
Deep understanding of distributed systems, microservices architecture, and containerization (Docker, Kubernetes)
Experience with monitoring and observability tools (Splunk, Datadog, Prometheus, Grafana)
Strong knowledge of networking, load balancing, CDN technologies, and DNS management
Proficiency in scripting languages (Python, Bash, PowerShell, etc.)
Soft Skills
Exceptional communication skills with ability to articulate complex technical issues to both technical and non-technical stakeholders
Executive presence and ability to effectively communicate with senior leadership during high-pressure incidents and post-mortems
Strong analytical and problem-solving abilities with a systematic approach to troubleshooting
Ability to remain calm under pressure and make critical decisions during incidents
Excellent collaboration skills with experience working across global, cross-functional teams
Strong documentation skills and attention to detail
Subsidiary:
PayPalTravel Percent:
0PayPal does not charge candidates any fees for courses, applications, resume reviews, interviews, background checks, or onboarding. When making an application directly, we will never ask you to share passwords, one-time passcodes (OTP), or verification codes. Any such request is a red flag and likely part of a scam. All communication regarding your application will come from official PayPal email domains. If you suspect fraudulent activity, please report it immediately. To learn more about how to identify and avoid recruitment fraud please visit https://careers.pypl.com/contact-us.
For the majority of employees, PayPal's balanced hybrid work model offers 3 days in the office for effective in-person collaboration and 2 days at your choice of either the PayPal office or your home workspace, ensuring that you equally have the benefits and conveniences of both locations.
Our Benefits:
At PayPal, we’re committed to building an equitable and inclusive global economy. And we can’t do this without our most important asset-you. That’s why we offer comprehensive, choice-based programs, to support all aspects of personal wellbeing—physical, emotional, and financial—delivering meaningful value where it matters most. We strive to create a flexible, balanced work culture with a holistic approach to benefits, including generous paid time off, healthcare coverage for you and your family, and resources to create financial security and support your mental health.
Who We Are:
Click Here to learn more about our culture and community.
Commitment to Diversity and Inclusion
PayPal provides equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, pregnancy, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state, or local law. In addition, PayPal will provide reasonable accommodations for qualified individuals with disabilities. If you are unable to submit an application because of incompatible assistive technology or a disability, please contact us at [email protected].
Belonging at PayPal:
Our employees are central to advancing our mission, and we strive to create an environment where everyone can do their best work with a sense of purpose and belonging. Belonging at PayPal means creating a workplace with a sense of acceptance and security where all employees feel included and valued. We are proud to have a diverse workforce reflective of the merchants, consumers, and communities that we serve, and we continue to take tangible actions to cultivate inclusivity and belonging at PayPal.
Any general requests for consideration of your skills, please Join our Talent Community.
We know the confidence gap and imposter syndrome can get in the way of meeting spectacular candidates. Please don’t hesitate to apply.
Skills Required
- 3+ years of relevant experience and a bachelor's degree, or an equivalent combination of education and experience
- 5+ years of experience in site reliability engineering, infrastructure operations, or similar technical operations roles
- Significant hands-on experience with AWS or GCP; multi-cloud experience preferred
- Professional-level certification in a major cloud platform, such as AWS Solutions Architect Professional or Google Cloud Professional Cloud Architect
- Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, Pulumi, or equivalent
- Significant hands-on experience with Kubernetes and CNCF ecosystem tools
- Experience managing critical incidents in Infrastructure-as-Code-driven environments, including IaC state issues, GitOps failures, and cloud-native deployment problems
- Ability to read and review Python, Go, Bash, YAML, HCL, and JSON during incident troubleshooting
- Experience with monitoring and observability tools including Splunk, Datadog, Prometheus, or Grafana
- Strong knowledge of networking, load balancing, CDN technologies, and DNS management
- Proficiency in operational automation scripting using Python, Bash, PowerShell, or equivalent
- Deep understanding of distributed systems, microservices architecture, and containerization
- Exceptional communication skills and executive presence during high-pressure incidents and postmortems
- Strong analytical, problem-solving, collaboration, documentation, and decision-making skills
PayPal Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about PayPal and has not been reviewed or approved by PayPal.
-
Healthcare Strength — Health coverage starts on the date of hire with medical, dental, vision, wellness resources, and health-navigation support. Eligibility includes spouses/domestic partners and dependents up to age 26, indicating broad and robust coverage.
-
Leave & Time Off Breadth — Flexible time-off frameworks and a market-leading sabbatical after five years provide substantial time-away options. Paid leaves span bonding/parental, disability, and localized programs, offering broad coverage across situations.
-
Retirement Support — A 401(k) with company match and a year-end true-up strengthens long-term savings. Financial-wellbeing tools and related programs further support retirement planning.
PayPal Insights
What We Do
HELP US REIMAGINE MONEY. At PayPal, we believe that now is the time to democratize financial services so that moving and managing money is a right for all citizens, not just the affluent. We are driven by this purpose, and we uphold our cultural values of collaboration, innovation, wellness and inclusion as our guide for making decisions and conducting business every day. It is our duty and privilege to be customer champions and put those we serve at the center of everything we do. We are one team that respects and values diversity of thought for everyone, everywhere, and we actively seek to create an energizing workplace that brings out the best in all of us. If you’re ready to shape the future of money, join the team at PayPal. We're proud to work here. You will be too. PayPal is headquartered in San Jose, California and its international headquarters is located in Singapore.








