Site Reliability Engineering (SRE) Lead

Posted Yesterday
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Artificial Intelligence • Cloud • Information Technology • Automation
The Role
Lead SRE responsible for platform reliability, incident management, automation, observability, and production stability across cloud-native and Kubernetes environments. Mentor teams, run RCA, define SLIs/SLOs, automate operations, manage Apigee APIs, and support CI/CD and IaC initiatives to improve availability and performance.
Summary Generated by Built In

Job Title: Site Reliability Engineering (SRE) Lead


Location: Hyderabad / Mumbai
Employment Type: Full-Time

 

About SID Global Solutions

SID Global Solutions (SIDGS) is a leading Digital Engineering and Technology Services company specializing in Cloud, API Management, DevOps, Platform Engineering, Kubernetes, Microservices, and Digital Transformation. We are looking for an experienced Site Reliability Engineering (SRE) Lead to drive platform reliability, operational excellence, and production stability for mission-critical enterprise applications.

 

Job Summary

We are seeking a highly skilled SRE Lead with strong expertise in cloud infrastructure, Kubernetes, API Management, and production operations. The ideal candidate will be responsible for ensuring the availability, scalability, performance, and reliability of enterprise applications while leading incident management, observability, and automation initiatives.

In this role, you will work closely with Development, DevOps, Infrastructure, Platform Engineering, and Application Support teams to maintain highly available production environments, reduce operational risks, and improve service reliability through automation and proactive monitoring.

 

Key Responsibilities

Reliability Engineering

  • Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to improve platform reliability.
  • Review application architecture and infrastructure designs to ensure scalability, resilience, and high availability.
  • Drive reliability improvements across cloud-native applications and distributed systems.
  • Identify opportunities to eliminate operational bottlenecks through automation and process improvements.

Incident & Production Management

  • Lead Major Incident Management (P1/P2) activities and act as the primary technical escalation point for critical production issues.
  • Coordinate with cross-functional teams to restore services within defined SLAs.
  • Conduct Root Cause Analysis (RCA) and drive preventive and corrective actions to reduce recurring incidents.
  • Participate in Change Management and Release activities to ensure production stability.
  • Prepare incident reports and communicate status updates to business and technology stakeholders.

Automation & Platform Engineering

  • Develop automation scripts and self-healing solutions to improve operational efficiency.
  • Build reusable operational runbooks and standard operating procedures.
  • Automate routine operational tasks using Python, Bash, or similar scripting languages.
  • Support Infrastructure as Code (IaC) initiatives and CI/CD pipeline improvements.

Monitoring & Observability

  • Design and maintain enterprise monitoring and alerting solutions.
  • Create dashboards and alerts to proactively monitor application health, infrastructure, APIs, and Kubernetes environments.
  • Analyze performance trends and recommend improvements for system stability and capacity planning.
  • Ensure effective monitoring coverage across production environments.

API & Kubernetes Administration

  • Manage and troubleshoot Google Apigee API Gateway configurations, policies, and traffic routing.
  • Monitor Kubernetes clusters, workloads, namespaces, ingress controllers, and container health.
  • Optimize application performance and resource utilization within Kubernetes environments.
  • Support production deployments and post-release validation activities.

Leadership & Collaboration

  • Mentor SRE, DevOps, and Production Support engineers.
  • Establish operational best practices, troubleshooting guidelines, and technical documentation.
  • Collaborate with Development, QA, Infrastructure, Security, and Business teams to improve platform reliability.
  • Drive a culture of continuous improvement, automation, and operational excellence.

 

Required Technical Skills

  • Strong experience with Google Cloud Platform (GCP).
  • Hands-on expertise in Google Apigee API Management.
  • Experience managing production Kubernetes (K8s) environments.
  • Good understanding of NGINX, reverse proxy, and load balancing concepts.
  • Strong knowledge of Linux/Unix Administration.
  • Experience with Python, Bash, or Go scripting.
  • Familiarity with CI/CD pipelines, Git, and Jenkins.
  • Understanding of Infrastructure as Code (Terraform or equivalent is preferred).

Monitoring & Observability Tools

Experience with one or more of the following:

  • Datadog
  • Dynatrace
  • Prometheus
  • Grafana
  • Splunk
  • ELK Stack
  • AppDynamics

 

Required Qualifications

  • 7–10 years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, or Production Support.
  • Strong experience supporting enterprise production environments.
  • Hands-on experience with Kubernetes, GCP, and Google Apigee.
  • Good understanding of distributed systems, microservices architecture, and cloud-native applications.
  • Experience in Incident, Problem, Change, and Release Management.
  • Ability to troubleshoot complex production issues and coordinate cross-functional teams during critical incidents.

 

Preferred Qualifications

  • Experience in Banking, Financial Services, or other enterprise environments.
  • ITIL Foundation certification.
  • Google Cloud Professional Certification.
  • Kubernetes Certification (CKA/CKAD) is an added advantage.
  • Experience with Service Mesh technologies such as Istio is preferred.

 



Skills Required

  • 7-10 years experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Support
  • Hands-on experience with Google Cloud Platform (GCP)
  • Hands-on experience with Google Apigee API Management
  • Experience managing production Kubernetes (K8s) environments
  • NGINX, reverse proxy, and load balancing knowledge
  • Linux/Unix administration experience
  • Scripting with Python, Bash, or Go
  • Familiarity with CI/CD pipelines, Git, and Jenkins
  • Infrastructure as Code (Terraform or equivalent)
  • Experience with monitoring and observability tools (Datadog, Dynatrace, Prometheus, Grafana, Splunk, ELK, AppDynamics)
  • Experience in Incident, Problem, Change, and Release Management and leading major incident response (P1/P2)
  • Strong understanding of distributed systems, microservices architecture, and cloud-native applications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
339 Employees
Year Founded: 2006

What We Do

SID Global Solutions (SIDGS) is an AI-first digital transformation and consulting company serving enterprises worldwide. It delivers intelligent full-stack solutions across artificial intelligence, cloud modernization, automation, API and application transformation, and enterprise intelligence. SIDGS helps organizations modernize legacy systems, optimize operations, integrate data and platforms, and accelerate measurable digital growth through scalable, customer-centric technology services for complex and demanding enterprise environments.

Similar Jobs

Micron Technology Logo Micron Technology

Principal Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Hyderabad, Telangana, IND
45000 Employees

Micron Technology Logo Micron Technology

Senior Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Hyderabad, Telangana, IND
45000 Employees

Vertafore Logo Vertafore

Technical Lead

Information Technology • Insurance • Software
Hybrid
Hyderabad, Telangana, IND
2372 Employees

Nasuni Logo Nasuni

Principal Search Engineer - Java, Elastic Search, Solr

Artificial Intelligence • Big Data • Cloud • Security • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Easy Apply
Hybrid
Hyderabad, Telangana, IND
550 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account