Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
1M-2M Annually
Senior level
Artificial Intelligence • HR Tech • Professional Services • Software
The Role
Build, operate, and improve reliable production systems across hybrid and multi-cloud environments. Manage Kubernetes, AWS, IBM Cloud, Terraform infrastructure, containerized workloads, observability platforms, networking, disaster recovery, and automation. Define reliability objectives, lead high-severity incident response, conduct root-cause analysis, improve system resilience, and support capacity planning. The role also involves architecture collaboration, compliance-focused infrastructure practices, mentoring engineers, and promoting SRE and DevOps standards.
Summary Generated by Built In

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟭𝟯𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟮𝟬𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟭𝟯-𝟮𝟬 𝗟𝗣𝗔)

Experience: 4+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are looking for an experienced Senior Site Reliability Engineer (SRE) to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems across hybrid and multi-cloud environments.

The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms, along with a strong understanding of production operations and distributed systems.


Requirements

Key Responsibilities

  • Define and manage SLIs, SLOs, SLAs, error budgets, and reliability objectives for critical production services.
  • Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
  • Manage and support production Kubernetes environments, including Amazon EKS and Red Hat OpenShift.
  • Deploy and maintain containerised workloads using Docker, Kubernetes, and Helm.
  • Manage cloud infrastructure across AWS and IBM Cloud, including hybrid-cloud environments.
  • Design and maintain reliable cloud connectivity, networking, disaster-recovery, and failover solutions.
  • Develop and maintain infrastructure using Terraform and Infrastructure as Code (IaC) practices.
  • Automate operational processes, infrastructure tasks, and troubleshooting workflows using Python and Bash.
  • Build and enhance observability solutions using Prometheus, Grafana, OpenTelemetry, Thanos, and logging platforms.
  • Monitor system health, identify performance bottlenecks, and proactively address reliability risks.
  • Participate in and lead high-severity incident response and production troubleshooting.
  • Conduct root-cause analysis and lead post-incident reviews and corrective actions.
  • Develop and maintain capacity-planning and reliability-improvement strategies.
  • Implement secure, resilient, and compliant infrastructure practices across cloud environments.
  • Support disaster-recovery planning, testing, and continuous improvement.
  • Collaborate with software engineering, platform, security, and architecture teams to improve production reliability.
  • Contribute to architecture reviews, engineering standards, operational best practices, and automation initiatives.
  • Mentor engineers and promote strong SRE, DevOps, observability, and production-engineering practices.

What Makes You a Great Fit

  • 4–6 years of professional experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or a closely related field.
  • Strong hands-on experience with AWS and Kubernetes in production environments.
  • Experience managing Amazon EKS, Docker, and Helm.
  • Practical experience with Red Hat OpenShift is highly desirable.
  • Strong proficiency in Terraform and Infrastructure as Code practices.
  • Hands-on scripting and automation experience using Python and Bash.
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Experience with OpenTelemetry, Thanos, logging platforms, or similar observability technologies.
  • Strong understanding of SLIs, SLOs, error budgets, incident management, and production troubleshooting.
  • Good understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity.
  • Experience working with hybrid or multi-cloud infrastructure, preferably including AWS and IBM Cloud.
  • Strong understanding of containers, distributed systems, scalability, availability, and fault tolerance.
  • Experience with disaster recovery, capacity planning, and production resilience.
  • Exposure to regulated or compliance-driven environments such as HIPAA, SOC 2, PCI DSS, or ISO 27001.
  • Strong analytical, troubleshooting, and root-cause analysis skills.
  • Excellent communication and collaboration skills.
  • Ability to take ownership of critical production systems and operate effectively during high-severity incidents.
  • Experience mentoring engineers and contributing to technical architecture and reliability standards.

Skills Required

  • 4-6 years of professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, or a related field
  • Hands-on production experience with AWS and Kubernetes
  • Experience managing Amazon EKS, Docker, and Helm
  • Experience with Red Hat OpenShift
  • Strong proficiency with Terraform and Infrastructure as Code practices
  • Hands-on scripting and automation experience using Python and Bash
  • Strong experience with Prometheus and Grafana for monitoring and observability
  • Experience with OpenTelemetry, Thanos, logging platforms, or similar observability technologies
  • Understanding of SLIs, SLOs, error budgets, incident management, and production troubleshooting
  • Understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity
  • Experience with hybrid or multi-cloud infrastructure, preferably AWS and IBM Cloud
  • Understanding of containers, distributed systems, scalability, availability, and fault tolerance
  • Experience with disaster recovery, capacity planning, and production resilience
  • Exposure to regulated or compliance-driven environments such as HIPAA, SOC 2, PCI DSS, or ISO 27001
  • Strong analytical, troubleshooting, and root-cause analysis skills
  • Excellent communication and collaboration skills
  • Ability to own critical production systems and operate during high-severity incidents
  • Experience mentoring engineers and contributing to technical architecture and reliability standards
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2021

What We Do

Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

Similar Jobs

Zscaler Logo Zscaler

Site Reliability Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Hybrid
Bangalore, Bengaluru, Karnataka, IND
8697 Employees

NVIDIA Logo NVIDIA

Senior Site Reliability Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees

2K Logo 2K

Senior Site Reliability Engineer

Gaming • Information Technology • Mobile • Software • Esports
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
4200 Employees

The Walt Disney Company Logo The Walt Disney Company

Senior Site Reliability Engineer

Digital Media • Gaming • News + Entertainment • Sports
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
219548 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account