Sr SRE Engineer

Posted 9 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Information Technology • Consulting
The Role
Lead site reliability engineering for cloud-native platforms by defining SLIs and SLOs, managing error budgets, leading production incident response and root cause analysis, maintaining runbooks, and improving observability, automation, availability, and resilience. The role uses GCP, Kubernetes, Terraform, Helm, Prometheus, Grafana, ELK, and incident management tools while supporting capacity planning, failover testing, and cross-functional reliability reviews.
Summary Generated by Built In
Experience Required:
6-8 years of SRE or infrastructure engineering experience in cloud-native environments.

Mandatory:
• Cloud: GCP (GKE, Load Balancing, VPN, IAM)
• Observability: Prometheus, Grafana, ELK, Datadog
• Containers & Orchestration: Kubernetes, Docker
• Incident Management: On-call, RCA, SLIs/SLOs
• IaC: Terraform, Helm
• Incident Tools: PagerDuty, OpsGenie

Nice to Have:
• GCP Monitoring, Skywalking
• Service Mesh, API Gateway
• GCP Spanner, MongoDB (basic)

Scope:
• Drive operational excellence and platform resilience
• Reduce MTTR, increase service availability
• Own incident and RCA processes

Roles and Responsibilities:

•Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and manage error budgets across services.
• Lead incident management for critical production issues – drive root cause analysis (RCA) and postmortems.
• Create and maintain runbooks and standard operating procedures for high availability services.
• Design and implement observability frameworks using ELK, Prometheus, and Grafana; drive telemetry adoption.
• Coordinate cross-functional war-room sessions during major incidents and maintain response logs.
• Develop and improve automated system recovery, alert suppression, and escalation logic.
• Use GCP tools like GKE, Cloud Monitoring, and Cloud Armor to improve performance and security posture.
• Collaborate with DevOps and Infrastructure teams to build highly available and scalable systems.
• Analyze performance metrics and conduct regular reliability reviews with engineering leads.
• Participate in capacity planning, failover testing, and resilience architecture reviews.

Skills Required

  • 6-8 years of SRE or infrastructure engineering experience in cloud-native environments
  • Experience with GCP, including GKE, Load Balancing, VPN, and IAM
  • Experience with Prometheus, Grafana, ELK, and Datadog
  • Experience with Kubernetes and Docker
  • Experience with on-call operations, root cause analysis, and SLIs/SLOs
  • Experience with Terraform and Helm
  • Experience with PagerDuty and OpsGenie
  • Experience with GCP Monitoring and SkyWalking
  • Experience with service mesh and API Gateway
  • Basic experience with GCP Spanner and MongoDB
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Vaughan, Ontario
345 Employees
Year Founded: 2007

What We Do

@TechBlocks we power the software defined industries (SDI) of today and tomorrow. We are a software engineering and consulting firm. We build modern digital value chains and businesses reimagined to create frictionless experiences for innovative monetization methods and drive unforeseen efficiencies. We are known to build world class custom platforms and products that are cloud native for some of the worlds largest brands. We are the go to technology partners for born in digital businesses that grew with us from "Concept to Commercialization" and have revenues between $100M - $10B. We help modern businesses transition just from a technology outsourcing mentality to help create globally distributed digital COEs and mature them. Our converged COEs that we create in partnership with our clients help power software factories that are extremely dynamic. We have created modern digital COEs and factories that are created with a single minded goal to future proof our clients businesses. Everything we do is centred around two philosophies and practices - Design Thinking and Lean Engineering. Whether it is building digital commerce platforms, marketplace for worlds largest retailers or smart utilities applications and products or digital health products/platforms that power wearables, patches or devices across healthcare landscape; we do it all with speed and sophistication that is unmatched in the industry

Similar Jobs

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

MetLife Logo MetLife

Senior Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees

Crunchyroll Logo Crunchyroll

Senior Site Reliability Engineer

Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Remote or Hybrid
Hyderabad, Telangana, IND
1300 Employees

ModMed Logo ModMed

Senior Site Reliability Engineer

Healthtech • Software • Telehealth
In-Office
Hyderabad, Telangana, IND
1153 Employees

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account