Lead Site Reliability Engineer

Posted 13 Days Ago
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Digital Media • Gaming • News + Entertainment • Sports
The Role
Leads reliability, scalability, security, and performance for large-scale Disney Experiences platforms. Architects cloud-native and hybrid infrastructure, develops CI/CD and Infrastructure as Code automation, establishes observability and SLOs, leads incident response, and drives continuous improvement. The role mentors engineers, coordinates platform projects, collaborates across technical teams, and applies FinOps principles to optimize cloud costs and operational efficiency.
Summary Generated by Built In

Job Posting Title:

Lead Site Reliability Engineer

Req ID:

10159518

Job Description:

The Lead Site Reliability Engineer is a seasoned subject matter expert who drives reliability, scalability, and performance for critical Disney Experiences platforms that power immersive guest interactions across theme parks, resorts, cruise, vacation, travel, retail, and consumer experiences. In this lead role, you will guide other engineers, set technical direction for complex systems, and ensure our digital and physical experiences remain highly available, secure, and resilient for guests around the world.

Responsibilities:

As a Lead Site Reliability Engineer, you will serve as a skilled, experienced problem-solver and team player, acting as the "go-to" technical lead for assigned platforms and services across Disney Experiences. You will architect and evolve cloud-native and hybrid infrastructures, champion observability and DevOps practices, and mentor individual contributors to deliver reliable, secure, and cost‑effective solutions that support Disney's creative, customer-focused, and innovative guest experiences. This role matters because it safeguards the technology behind our stories, requiring deep technical expertise, strong collaboration, and the ability to explain complex concepts and influence diverse stakeholders.

  • Architect, design, and build scalable, maintainable, and secure infrastructure and platforms, including cloud-native and container-based solutions, to support mission-critical commerce and guest-facing applications.
  • Lead the evolution of DevOps and SRE practices by consulting on, designing, and supporting CI/CD pipelines, automating infrastructure and operations, and creating telemetry and observability for monitoring and incident response.
  • Serve as the SRE subject matter expert and technical lead for assigned products and platforms, owning reliability strategies, defining SLIs/SLOs/SLAs, and driving continuous improvement in uptime and performance.
  • Identify root causes of operational issues in large-scale distributed systems, lead major incident response, and deliver clear retrospectives and remediation plans that reduce future risk and operational toil.
  • Develop, maintain, and enhance automation, scripts, and Infrastructure as Code to standardize deployments, improve reliability, and support complex, non-standard environments without relying solely on runbooks.
  • Collaborate with product, engineering, security, and operations teams to plan capacity, monitoring, configuration, security, metrics, reporting, recovery, and migration strategies for new initiatives and events impacting supported platforms.
  • Mentor, train, and guide other engineers by providing continuous coaching, feedback, and technical direction, holding self and others accountable to commitments and aligning team work with organizational goals.
  • Plan and coordinate team efforts and platform-oriented projects with moderate complexity and risk, breaking down organizational goals into clear outcomes and negotiating solutions to complex reliability challenges.
  • Apply FinOps and cost-optimization principles to cloud environments, implementing governance, tagging, rightsizing, and usage analysis to balance reliability, performance, and cost efficiency.
  • Champion a diverse, inclusive, team-oriented culture that encourages innovation, creative problem solving, and service-minded collaboration, ensuring every voice is heard and Disney values are experienced daily.

Required Qualifications:

  • Minimum 7 years of related work experience in Site Reliability Engineering, Systems Engineering, or software development, with a focus on large-scale, distributed, and cloud-based systems.
  • Bachelor's Degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent work experience.
  • Extensive hands-on experience with cloud hosting services (AWS, Azure, Google Cloud) and modern cloud architectures, including containers and orchestration platforms such as Docker, Kubernetes, ECS, AKS, and GKE.
  • Proficiency in Infrastructure as Code and configuration management tools (e.g., Terraform, CloudFormation, Ansible, Chef) and CI/CD pipelines using tools such as GitHub, GitLab, Jenkins, AWS CodeBuild, or Azure DevOps.
  • Fluency in core scripting and programming languages (e.g., Python, NodeJS, Golang, Bash, Perl, Ruby, Java) and strong UNIX/Linux administration, troubleshooting, and security skills.
  • Applied expertise in observability and monitoring, including defining and implementing SLIs, SLOs, SLAs and using major APM and logging tools (e.g., AppDynamics, New Relic, ELK stack, Datadog, Splunk, New Relic).
  • Strong knowledge of networking and distributed systems, including HTTP, TCP/IP, DNS, TLS, SSH, VPCs, gateways, firewalls, and microservices architectures.
  • Experience with databases and data platforms such as MySQL, MongoDB, DynamoDB, Redis, and data solutions like Snowflake or Tableau, including ELT processes for data-driven decision making.
  • Demonstrated ability to lead technical projects, evaluate new systems and infrastructure solutions for feasibility, and design reliable, scalable enterprise systems in agile environments.
  • Outstanding troubleshooting methodology and communication skills, with the ability to explain difficult concepts, influence without direct authority, and mentor, train, and guide other engineers.
  • Ability to plan and prioritize work aligned with organizational goals, collaborate effectively across teams, and hold self and others accountable to meet commitments.

Preferred Qualifications:

  • Experience leveraging AI and automation for predictive insights and continuous improvement in system reliability and operational efficiency.
  • Expertise in cloud infrastructure design and dynamic development technologies using Java, NodeJS, Python, and relational databases in large-scale business environments.
  • Background operating production container environments and multi-origin hybrid (cloud and on-premise) architectures.
  • Master's degree in Computer Science, Information Systems, or a related technical discipline.

 

Job Posting Segment:

DX Technology

Job Posting Primary Business:

Tech Delivery, Platforms, & Core Systems

Primary Job Posting Category:

Site/System Reliability Engineer

Employment Type:

Full time

Primary City, State, Region, Postal Code:

Bangalore, India

Alternate City, State, Region, Postal Code:

Date Posted:

2026-09-03

Skills Required

  • Minimum 7 years of related experience in Site Reliability Engineering, Systems Engineering, or software development focused on large-scale, distributed, cloud-based systems.
  • Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent work experience.
  • Hands-on experience with AWS, Azure, or Google Cloud and cloud architectures using Docker, Kubernetes, ECS, AKS, or GKE.
  • Proficiency with Infrastructure as Code and configuration management tools such as Terraform, CloudFormation, Ansible, or Chef.
  • Experience building and supporting CI/CD pipelines using GitHub, GitLab, Jenkins, AWS CodeBuild, or Azure DevOps.
  • Fluency in scripting or programming languages including Python, Node.js, Golang, Bash, Perl, Ruby, or Java.
  • Strong Unix/Linux administration, troubleshooting, and security skills.
  • Expertise in observability, monitoring, SLIs, SLOs, SLAs, APM, and logging tools such as AppDynamics, New Relic, ELK Stack, Datadog, or Splunk.
  • Strong knowledge of networking and distributed systems, including HTTP, TCP/IP, DNS, TLS, SSH, VPCs, gateways, firewalls, and microservices.
  • Experience with MySQL, MongoDB, DynamoDB, Redis, Snowflake, or Tableau, including ELT processes.
  • Ability to lead technical projects and design reliable, scalable enterprise systems in agile environments.
  • Strong troubleshooting, communication, stakeholder influence, mentoring, training, planning, and prioritization skills.
  • Experience using AI and automation for predictive reliability insights and operational improvement.
  • Experience with Java, Node.js, Python, relational databases, production container environments, and hybrid cloud/on-premise architectures.
  • Master’s degree in Computer Science, Information Systems, or a related technical discipline.

The Walt Disney Company Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about The Walt Disney Company and has not been reviewed or approved by The Walt Disney Company.

  • Pay Growth & Progression Recent union agreements raised wage floors for large groups of park cast members—e.g., Disneyland’s $24/hour minimum rising to $26 over the contract and Walt Disney World’s path from $18 toward about $20–$20.50 by 2026—signaling upward movement in hourly pay. These steps are described as meaningful improvements for many frontline roles.
  • Healthcare Strength Company materials outline medical, dental, and vision coverage for many full‑time roles, wellness resources, and (in Central Florida) access to Centers for Living Well clinics and pharmacy. References to mental‑health support and paid time off reinforce a strong core health offering.
  • Wellbeing & Lifestyle Benefits Complimentary theme‑park admission and discounts on hotels, dining, merchandise, and recreation are positioned as signature perks. Education support through Disney Aspire adds notable lifestyle value for eligible hourly employees.

The Walt Disney Company Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Burbank, CA
219,548 Employees
Year Founded: 1923

What We Do

The Walt Disney Company is a leading diversified international family entertainment and media enterprise that operates through segments including entertainment, sports, and experiences.

Similar Jobs

Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Zeta Global Logo Zeta Global

Site Reliability Engineer

AdTech • Artificial Intelligence • Marketing Tech • Software • Analytics
Easy Apply
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
2429 Employees
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
52655 Employees

Similar Companies Hiring

Bankrate Thumbnail
Artificial Intelligence • Consumer Web • Digital Media • Fintech • Marketing Tech • Software • Financial Services
US
160 Employees
ARB Interactive Thumbnail
Gaming • Mobile • Software
Miami, Florida
190 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account