Site Reliability Engineer - Warehousing IT Operations

Posted Yesterday
Be an Early Applicant
Taguig City, Fourth District NCR, National Capital Region, PHL
In-Office
Senior level
AdTech • Beauty • Marketing Tech • Retail • Pharmaceutical
The Role
Lead incident response and reliability efforts for warehousing IT systems. Design, implement, and automate monitoring, alerting, and resilient architecture. Troubleshoot and perform root cause analysis, collaborate with engineering/DevOps, mentor team members, and manage on-call rotations to ensure high availability and performance.
Summary Generated by Built In

Job Location

Taguig City

Job Description

As a Site Reliability Engineer (SRE) in the Warehousing IT Operations – Incident Response Team, you will be responsible for leading incident response efforts, ensuring swift and effective resolution of critical system issues. You will also play a critical role in ensuring the reliability, scalability, and performance of our systems and services. SRE combines software engineering and operations to build, maintain, and support highly available and efficient infrastructure. Your expertise in troubleshooting and root cause analysis will be essential in identifying and addressing the underlying causes of incidents. You will work closely with software engineers, DevOps teams, and other stakeholders to implement preventive measures and enhance system resilience. Collaborating with cross-functional teams, you will design, implement, and automate robust systems, monitoring tools, and processes. With a strong focus on stability and uptime, you will proactively identify and resolve performance bottlenecks, optimize system architecture, and drive continuous improvement. Your keen eye for continuous improvement will also drive post-incident reviews and contribute to the creation of incident management best practices. By actively monitoring system health, responding to incidents in a timely manner, and implementing proactive measures, you will play a pivotal role in maintaining the stability and availability of our services, ensuring an exceptional user experience for our customers.

This is a Managerial position. Being a manager at P&G involves leading teams and / or end-to-end processes, managing P&G resources, and driving business results. Managers are responsible for overseeing various aspects of the business, including strategy, operations, and team performance. They play a crucial role in ensuring that P&G's brands continue to grow and succeed in the market. Managers at P&G are expected to have strong leadership skills, a growth mindset, and the ability to make data-driven decisions roles lead and initiatives, significantly impacting business results through independent judgment and minimal guidance.

How success looks like

Success as a Site Reliability Engineer (SRE) involves different areas of the role including incident response, monitoring and reliability, and effectively collaborating with customers and users, addressing their needs and expectations:

  • Incident Response: Swiftly respond to and resolve critical incidents, ensuring minimal impact on system availability and user experience while driving continuous improvement in incident management processes.
  • Reliability: Ensure high system availability and reliability through robust monitoring, optimization of system architecture, and cross-functional collaboration to design and implement resilient systems.
  • Monitoring: Implement comprehensive monitoring solutions to gain real-time insights into system performance, enabling proactive incident response and continuous improvement of system visibility and resource optimization.
  • Working with Customers/Users: Collaborate directly with customers and users to understand their needs, proactively address concerns, and provide exceptional customer support to ensure reliable and performant systems that meet their expectations.

Responsibilities: 

Incident Response:

  • Lead incident response efforts, swiftly resolving critical incidents to minimize downtime and user impact.
  • Implement effective incident management processes, ensuring clear communication, coordination, and documentation.
  • Conduct root cause analysis, implementing preventive measures and driving continuous improvement.

Reliability:

  • Ensure high system availability through robust monitoring, alerting, and automated incident response systems.
  • Optimize system architecture and configurations for improved performance, scalability, and fault tolerance.
  • Collaborate cross-functionally to design and implement resilient systems using industry best practices.
  • Implement comprehensive monitoring solutions, providing real-time insights into system performance and health.
  • Configure and manage monitoring tools, ensuring accurate and actionable alerts for proactive incident response.
  • Continuously evaluate and enhance monitoring strategies to improve system visibility and resource optimization.

Upskilling:

  • Stay updated with industry trends, technologies, and best practices in Site Reliability Engineering.
  • Continuously develop technical skills in system architecture, automation, cloud technologies, and incident response.
  • Share knowledge, mentor team members, and foster a culture of learning and upskilling.

Managing Users/Customers’ Needs and Expectations:

  • Collaborate directly with users and customers to understand their needs and pain points.
  • Proactively address customer/user concerns, ensuring reliable and performant systems.
  • Provide exceptional customer support, communicate updates, resolutions, and gather feedback for continuous improvement.

Monitoring:

  • Implement comprehensive monitoring solutions, providing real-time insights into system performance and health.
  • Configure and manage monitoring tools, ensuring accurate and actionable alerts for proactive incident response.
  • Continuously evaluate and enhance monitoring strategies to improve system visibility and resource optimization.

Upskilling:

  • Stay updated with industry trends, technologies, and best practices in Site Reliability Engineering.
  • Continuously develop technical skills in system architecture, automation, cloud technologies, and incident response.
  • Share knowledge, mentor team members, and foster a culture of learning and upskilling.

Managing Users/Customers’ Needs and Expectations:

  • Collaborate directly with users and customers to understand their needs and pain points.
  • Proactively address customer/user concerns, ensuring reliable and performant systems.
  • Provide exceptional customer support, communicate updates, resolutions, and gather feedback for continuous improvement.

Job Qualifications

Technical Expertise and Experience:

  • Knowledge or familiarity in system administration, including Linux/Unix environments, cloud platforms (such as AWS, Azure, or GCP) and SAP.
  • Experience with configuration management tools and infrastructure-as-code frameworks (e.g., Terraform).
  • Proficiency in at least one programming language (e.g., Python, C#) and experience with scripting for automation tasks.
  • Understanding of networking protocols, network infrastructures, load balancing, and DNS management.
  • Familiarity with containerization and orchestration technologies (e.g., Docker, Kubernetes).
  • Familiarity with databases and proficiency in writing SQL queries.
  • Experience or familiarity with monitoring and observability tools (e.g., Prometheus, Grafana).
  • Knowledge of incident response methodologies, root cause analysis, and implementing preventive measures.
  • Understanding of security best practices and experience with implementing secure systems.
  • Experience in Warehousing Management Systems (e.g. RTCIS, PrIME) or Warehousing Operations is a plus.

Soft Skills:

  • Strong problem-solving and troubleshooting skills, with an ability to analyze complex issues and devise effective solutions.
  • Excellent communication and collaboration skills to work effectively with cross-functional teams, stakeholders, and customers.
  • Ability to thrive in a fast-paced, dynamic environment, managing multiple priorities and adapting to changing circumstances.
  • Strong attention to detail and a commitment to delivering high-quality work.
  • Proactive and self-motivated, with a continuous learning mindset and a drive for staying updated with industry trends and technologies.
  • Strong teamwork and interpersonal skills, with an ability to build relationships and work effectively in a collaborative environment.
  • Ability to thrive under pressure and effectively manage incidents, ensuring timely resolutions and minimizing downtime.

This role requires a commitment to work a standard 5-day workweek, with rotating on-call assignment on weekends (Saturday/Sunday) and holidays. The nature of the Site Reliability Engineer (SRE) position necessitates coverage and support across the week, ensuring the reliability and availability of our systems. This schedule allows for effective incident response and continuous monitoring of system health, as well as collaboration with cross-functional teams. We value work-life balance and will strive to provide a predictable and manageable schedule within this framework, while still meeting the needs of our customers and maintaining the stability of our services.

About us

We produce globally recognized brands and we grow the best business leaders in the industry. With a portfolio of trusted brands as diverse as ours, it is paramount our leaders are able to lead with courage the vast array of brands, categories and functions. We serve consumers around the world with one of the strongest portfolios of trusted, quality, leadership brands, including Always®, Ariel®, Gillette®, Head & Shoulders®, Herbal Essences®, Oral-B®, Pampers®, Pantene®, Tampax® and more. Our community includes operations in approximately 70 countries worldwide.

Visit http://www.pg.com to know more.

We are an equal opportunity employer and value diversity at our company. We do not discriminate against individuals on the basis of race, color, gender, age, national origin, religion, sexual orientation, gender identity or expression, marital status, citizenship, disability, HIV/AIDS status, or any other legally protected factor.

Job Schedule

Full time

Job Number

R000157076

Job Segmentation

Entry Level

Skills Required

  • Linux/Unix system administration experience
  • Experience with cloud platforms (AWS, Azure, or GCP)
  • Familiarity with SAP
  • Experience with infrastructure-as-code / Terraform
  • Proficiency in at least one programming language (Python or C#) and scripting for automation
  • Experience with configuration management tools
  • Understanding of networking protocols, load balancing, and DNS management
  • Familiarity with containerization and orchestration (Docker, Kubernetes)
  • Familiarity with databases and ability to write SQL queries
  • Experience with monitoring and observability tools (Prometheus, Grafana)
  • Knowledge of incident response methodologies and root cause analysis
  • Understanding and application of security best practices for systems
  • Experience leading teams or managerial experience
  • Ability to participate in rotating on-call assignments including weekends and holidays
  • Experience with Warehousing Management Systems (e.g., RTCIS, PrIME) or warehousing operations

Procter & Gamble Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Procter & Gamble and has not been reviewed or approved by Procter & Gamble.

  • Fair & Transparent Compensation Compensation is considered competitive and benchmarked against top industry peers, supported by formal pay‑equity audits and stated transparent principles. Feedback suggests pay is a strong draw, with raises described as attainable and benefits starting from day one of training.
  • Healthcare Strength Health coverage is broad and immediate, including medical, dental, vision, life and disability insurance, with mental‑health and telemedicine offerings expanded recently. This breadth and day‑one access contribute to a perception of reliable, comprehensive care.
  • Parental & Family Support Parental leave is inclusive under a global framework for all parents, with added recovery time for birth mothers. Adoption, fertility, childcare support and eldercare services further strengthen family support.

Procter & Gamble Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Cincinnati, OH
117,512 Employees

What We Do

Procter & Gamble Company is an American multi-national consumer goods corporation. P&G was founded over 180 years ago as a soap and candle company. Today, we’re the world’s largest consumer goods company and home to iconic, trusted brands, including Always®, Charmin®, Braun®, Fairy®, Febreze®, Gillette®, Head & Shoulders®, Oral B®, Pantene®, Pampers®, Tide®, and Vicks®. The design, development, growth and success of these products—and many more—is thanks to the innovative and insightful minds of our people. From Day 1, you’ll help make everyday life easier for our 5 billion consumers through billion dollar brands. With our large global footprint, there are many opportunities to work with P&G in multiple locations. We offer opportunities in approximately 70 countries and continually aim to attract, reward and advance the finest people in the world. As a "build from within"​ organization, we see 95% of our people start at an entry level and progress through the organization. Here, we want you to get your career off to a fast start. That's why we don't have any rotational development programs or gradual ramping-up periods: you’ll be able—and encouraged—to dive right in from day 1. Join us and help make life better through meaningful work that makes an impact from Day 1.

Similar Jobs

Remitly Logo Remitly

Head of Operations (Senior Manager, Concierge Experience)

eCommerce • Fintech • Payments • Software • Financial Services
In-Office
Manila, Metro Manila, National Capital Region, PHL
2800 Employees

Remitly Logo Remitly

Accountant

eCommerce • Fintech • Payments • Software • Financial Services
In-Office
2 Locations
2800 Employees

Mondelēz International Logo Mondelēz International

Analytics Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
4 Locations
90000 Employees

Mondelēz International Logo Mondelēz International

Compensation Analyst

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
Manila, Metro Manila, National Capital Region, PHL
90000 Employees

Similar Companies Hiring

PRIMA Thumbnail
Travel • Software • Marketing Tech • Hospitality • eCommerce
US
15 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account