Senior Site Reliability Engineer

Reposted 3 Days Ago
Be an Early Applicant
Chennai, Tamil Nadu, IND
Hybrid
Senior level
Big Data • Cloud • Logistics • Machine Learning • Retail
The Role
Owner of availability and reliability for high-scale checkout services: triage and resolve incidents, build alerting/monitoring, automate remediation, tune performance, run DR planning, streamline deployments, develop tooling and dashboards, participate in on-call rotations, and drive RCA and reliability improvements across cloud-native microservice platforms.
Summary Generated by Built In
Position Summary...

What you'll do...

About Team:

Transactional System provides core transactional systems to enable
segment and technology partners in creating wonderful omni experiences with speed and leverage. We are a highly motivated group of engineers, working in an agile group to solve sophisticated and high impact problems. This role is part of Cloud Powered Checkout team and will build the next generation multi-tenant, client agnostic, highly scalable, omnichannel checkout solution to seamlessly enable a frictionless customer checkout experience across all sales channels globally. We process millions of orders daily through our high-performance checkout services running in Edge and Cloud.

As a Site Reliability Engineer in the CPC Team, you will work with L2, Other dependent Applications, Platform team, DevOps and Engineering practitioners to proactively maintain mission-critical infrastructure, cloud platforms, microservices, tools, and processes that will ensure the highest levels of availability and reliability of CPC applications.

Our team works closely with our US stores and eCommerce business to better serve customers by empowering team members, stores, and merchants with technological innovation. From groceries and entertainment to sporting goods and crafts, Walmart U.S. offers an extensive selection that our customers value, whether they shop online at Walmart.com, through one of our mobile apps, or in-store. Focus areas include customers, stores and employees, in-store service, merchant tools, merchant data science, and search and personalization.

What you'll do:

  • Incident triage, Escalation and Resolution: Triage site-impacting production issues by quantifying impact, severity and urgency, analyzing systems for quick remediation, engaging the right teams for recovery [Reduce MTTE Mean Time to Engage], and focusing on immediate restoration [ Reduce MTTR Mean Time to Restore] of large-scale enterprise systems.
  • Alert, Monitoring, Log analysis: Detect and analyze monitoring graphs and alerts to identify systems causing production impacts with various tools like Grafana, Prometheus, MMS, Service Now, JIRA, Dynatrace, Splunk etc [Reduce MTTD Mean Time to Detect].
  • Enhance Alerting solutions: Design and implement JavaScript for the integration of alerting tool with service API endpoints with various tools like ServiceNow, Spotlight, Splunk, and xMatters. Requires knowledge of: Monitoring and alerting tools; Monitoring metrics and key performance indicators (for example, availability, MTBF, MTTR); SLIs and SLOs (for example, request latency, availability, error rates, saturation); Distributed tracing; Alerting logic. To demonstrate awareness of the metrics used to monitor software or system performance. Monitors current performance data to ensure adherence to defined SLOs and SLIs for simple applications/systems. Demonstrates awareness of the different types of alerts generated by the monitoring tools. Demonstrates awareness of infrastructure and application metrics.
  • Disaster Recovery Planning: Requires knowledge of: Disaster recovery procedures and processes; Enterprise disaster recovery systems. To work with business partners to identify and document critical applications.
  • Performance and Optimization : Requires knowledge of: Unix/Linux performance optimization tuning; Java/NodeJS/Tomcat/Apache tuning and optimization; Chaos tools to utilize established criteria (for example, probability of failure, frequency of failure) to measure site reliability. Monitors site reliability conditions and new reliability requirements.
  • Work on Product Enrichment ; Content Services projects at Walmart: Develop enterprise monitoring and utilize tooling software solutions such as Grafana, Splunk etc, to improve visibility, pro-actively detect issues and restore system availability.
  • Develop Tools and support: Design and develop solutions for widespread internal communications for cloud applications support or workflows for infrastructure availability issues with various internal applications with multiple programming languages like Java, JavaScript (React, Node JS), Python and Shell programming technologies like Prometheus, Database Query languages. Design and develop a UI tool to display Item Content Quality data on a dashboard using AngularJS, ReactJs, HTML5 ; CSS3 etc
  • To create and maintain Playbooks.
  • Steps to perform correct analysis on the issues and engage correct teams for CPC, Dependent downstream services and Platform teams.
  • To handle Deployments. Streamline the deployments process and handle the responsibility as a single team. Understand and explore Post validations and back out steps to make app more resilient.
  • Coordinate with platform teams for non-app releases like VM upgrades, DB Maintenance, and other component environment related tasks.
  • Participate in rotating on-call duties and work across different time zone with a multi-national team
  • Responsible for timely root cause analysis [RCA] of production issues.
  • Develop reusable tooling and processes to drive and improve customer experience and lower operational costs.
  • Understand DevOps Industry best practices
  • Help teams to build highly Observable and Resilient systems
  • Collaborate with developers to capture requirements and understanding pain points
  • Build reusable tools, library, dashboards which can be used across DevOps/SRE teams

What you'll bring:

  • Bachelors degree in Computer Science, Engineering or related discipline
  • 5+ years of hands-on related to Site Reliability Engineer, Operations ; Development experience with Java Script, Java, Restful services, Git, Maven, Jenkins, DevOps, Containerization, Docker, Kubernetes, Azure, Google cloud, Kafka, Azure Cosmos, Azure SQL, Mega cache CI/CD ,Prometheus, Grafana, Splunk etc.
  • Automation and Self-healing: Demonstrate knowledge of scripting and software development for automation and self-healing of multi-cloud environments. Help enhance existing solutions by developing automation with Docker, Kubernetes and working with DevOps and Engineering partners.
  • Excellent end to end technical understanding of core infrastructure, cloud services, platforms, and micro-services.
  • Ability to effectively triage be able to detect and determine symptom vs cause.
  • Identify and drive continuous improvement efforts to reduce waste (eliminate, automate or streamline).
  • Influence the design of system architecture and tactical solutions.
  • Familiar with log centric tooling. Produce time series data and reusable dashboards for use both during and post event.

About Walmart Global Tech
Imagine working in an environment where one line of code can make life easier for hundreds of millions of people.  That’s what we do at Walmart Global Tech. We’re a team of software engineers, data scientists, cybersecurity expert's and service professionals within the world’s leading retailer who make an epic impact and are at the forefront of the next retail disruption. People are why we innovate, and people power our innovations. We are people-led and tech-empowered.

We train our team in the skillsets of the future and bring in experts like you to help us grow. We have roles for those chasing their first opportunity as well as those looking for the opportunity that will define their career. Here, you can kickstart a great career in tech, gain new skills and experience for virtually every industry, or leverage your expertise to innovate at scale, impact millions and reimagine the future of retail.

Flexible, hybrid work

Walmart’s culture sets us apart, and we know being together helps us innovate, learn and grow great careers. This role is based in our [Bangalore/Chennai] office for daily work, with the flexibility for associates to manage their personal lives.

Benefits

Beyond our great compensation package, you can receive incentive awards for your performance. Other great perks include a host of best-in-class benefits maternity and parental leave, pto, health benefits, and much more.

Belonging

We aim to create a culture where every associate feels valued for who they are, rooted in respect for the individual. Our goal is to foster a sense of belonging, to create opportunities for all our associates, customers and suppliers, and to be a Walmart for everyone.

At Walmart, our vision is "everyone included." by fostering a workplace culture where everyone is—and feels—included, everyone wins. Our associates and customers reflect the makeup of all 19 countries where we operate. By making Walmart a welcoming place where all people feel like they belong, we’re able to engage associates, strengthen our business, improve our ability to serve customers, and support the communities where we operate.

Equal opportunity employer

Walmart, inc., is an equal opportunities employer – by choice. We believe we are best equipped to help our associates, customers and the communities we serve live better when we really know them. That means understanding, respecting and valuing unique styles, experiences, identities, ideas and opinions – while being inclusive of all people.

Minimum Qualifications...

Outlined below are the required minimum qualifications for this position. If none are listed, there are no minimum qualifications.

Option 1: Bachelor's degree in computer science, computer engineering, computer information systems, software engineering, or related area and 3 years’ experience in site reliability engineering, site and system administration, infrastructure management, or related area.
Option 2: 5 years’ experience in site reliability engineering, site and system administration, infrastructure management, or related area.

Preferred Qualifications...

Outlined below are the optional preferred qualifications for this position. If none are listed, there are no preferred qualifications.

5 years’ experience in experience in site reliability engineering, site and system administration, infrastructure management, or related area., Master's degree in site reliability engineering, site and system administration, infrastructure management, or related area and 1 year’ experience in experience in site reliability engineering, site and system administration, infrastructure management, or related area., SRE certification (for example, IBM Cloud Site Reliability Engineer).

Primary Location...Tower 1, Part Of 1st, 2nd To 4th Flrs, Intl Tech Park Radial Road (radial It Park Pvt Ltd), , India

Skills Required

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or related discipline (or equivalent experience)
  • 3 years' SRE/site and system administration/infrastructure management experience with Bachelor's OR 5 years' experience without degree
  • 5+ years hands-on experience in site reliability engineering, operations, or related roles
  • Development experience with JavaScript and Java
  • Experience with RESTful services, Git, Maven, Jenkins, and CI/CD pipelines
  • Containerization and orchestration: Docker and Kubernetes
  • Cloud experience: Azure and Google Cloud Platform (GCP); familiarity with Azure Cosmos and Azure SQL
  • Monitoring, alerting and observability tooling: Prometheus, Grafana, Splunk, Dynatrace, MMS, ServiceNow, xMatters, JIRA
  • Scripting and automation: Python and Shell scripting for automation and self-healing
  • Performance tuning on Unix/Linux and Java/NodeJS/Tomcat/Apache
  • Experience with Kafka and database query languages (SQL)
  • Experience building UI/dashboard components with React or AngularJS, HTML5, CSS3 (for internal tooling)
  • Incident management skills: on-call participation, triage, RCA, and cross-team coordination
  • Preferred: Master's degree or SRE certification (for example IBM Cloud Site Reliability Engineer)

Walmart Global Tech Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Walmart Global Tech and has not been reviewed or approved by Walmart Global Tech.

  • Strong & Reliable Incentives Compensation packages typically include base pay, performance bonuses, and stock awards, with access to an employee stock purchase plan. Feedback suggests annual bonus structures and stock grants are a consistent part of total compensation in many roles.
  • Retirement Support Retirement offerings include a company 401(k) match and stock purchase options that support long‑term savings. Feedback suggests these programs are a notable strength within the overall package.
  • Parental & Family Support Family benefits feature paid parental leave and adoption assistance, alongside tax‑advantaged accounts for healthcare spending. Feedback suggests these supports contribute meaningfully to perceived total rewards value.

Walmart Global Tech Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bentonville, AR
578,950 Employees

What We Do

Walmart has a long history of transforming retail and using technology to deliver innovations that improve how the world shops and empower our 2.2 million associates. It began with Sam Walton and continues today with Global Tech associates working together to power Walmart and lead the next retail disruption. We’re a high-performing, primarily virtual workforce that is human-led and tech-empowered. Our world-class software engineers, data scientists and engineers, cybersecurity professionals, product managers and business service professionals work with top talent on cutting-edge technologies that create unique and innovative experiences for our associates, customers and members across Walmart, Sam’s Club and Walmart International. At Walmart Global Tech, one line of code or bold idea can make life easier for hundreds of millions of people – talk about epic impact at a global scale.

Similar Jobs

Walmart Global Tech Logo Walmart Global Tech

Senior Site Reliability Engineer

Big Data • Cloud • Logistics • Machine Learning • Retail
Hybrid
Chennai, Tamil Nadu, IND
578950 Employees

NVIDIA Logo NVIDIA

Senior Site Reliability Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees

Akamai Technologies Logo Akamai Technologies

Senior Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Weekday, Inc. Logo Weekday, Inc.

Senior Site Reliability Engineer

Artificial Intelligence • HR Tech • Professional Services • Software
In-Office
Chennai, Tamil Nadu, IND
3M-5M Annually

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account