Site Reliability Engineer (SRE) / Observability Engineer

Posted 8 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
HR Tech • Professional Services • Consulting
The Role
Build and operate observability infrastructure using Grafana, Prometheus, Loki, and related tools. Define and track SLOs, SLIs, and error budgets; create dashboards and actionable alerts; instrument distributed applications; automate runbooks and remediation; lead blameless incident reviews; and improve on-call practices, reliability, and capacity planning across Kubernetes-based systems.
Summary Generated by Built In
About the opportunity
We are hiring on behalf of a well-established global IT consulting and implementation firm with offices across North America, Europe, and India (HITEC City, Hyderabad). The organisation delivers technology solutions across Cloud, DevOps, SAP, and AI for enterprise clients globally and has a strong people-first, learning-oriented culture.

Role overview
We are looking for a Site Reliability Engineer with a strong Observability specialisation to drive service reliability, reduce operational toil, and build best-in-class monitoring and alerting infrastructure. The ideal candidate brings deep Grafana expertise and will take ownership of SLO/SLA definition, distributed system visibility, and driving the shift from reactive to proactive operations.

Key responsibilities
• Define, track, and report on Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across platform services
• Build, maintain, and optimise observability infrastructure using Grafana, Prometheus, Loki, Tempo, and related open-source tooling
• Develop dashboards and alerting rules that provide actionable, low-noise insights for engineering and operations teams
• Lead blameless post-incident reviews (PIRs) and drive systemic reliability improvements from learnings
• Partner with engineering teams to instrument applications with distributed tracing, structured logging, and custom metrics
• Reduce operational toil through automation — scripting runbooks, auto-remediation workflows, and self-healing infrastructure
• Define on-call practices, escalation policies, and runbooks; contribute to a sustainable on-call culture
• Evaluate and implement new observability tooling as the stack evolves (e.g., OpenTelemetry, Jaeger, VictoriaMetrics)

Required skills & experience
• 8+ years of combined SRE / DevOps / Platform Engineering experience
• Strong hands-on expertise with Grafana — dashboards, alerting, data sources
• Proficiency in Prometheus — PromQL, exporters, alertmanager
• Experience with log aggregation using Loki, ELK stack, or equivalent
• Solid understanding of distributed systems principles, microservices architecture, and container orchestration (Kubernetes)
• Proficiency in Python, Go, or Bash for automation and tooling
• Strong analytical thinking for root cause analysis and capacity planning

Good to have
• Hands-on experience with OpenTelemetry instrumentation
• Exposure to Grafana OnCall, Grafana Incident, or PagerDuty for incident management
• Familiarity with eBPF-based observability tools (Cilium, Parca)
• Azure or AWS certifications

What's on offer
• End-to-end ownership of observability — not just maintaining dashboards
• Hybrid work flexibility from HITEC City, Hyderabad
• Exposure to global-scale distributed systems for international clients
• Certification reimbursement and structured learning pathways

Location: Hyderabad (Hybrid)
Experience: 8+ years
Employment type: Full-time
Specialisation: Observability – Grafana, Prometheus, Loki stack

Skills Required

  • 8+ years of combined SRE, DevOps, or Platform Engineering experience
  • Hands-on expertise with Grafana dashboards, alerting, and data sources
  • Proficiency with Prometheus, PromQL, exporters, and Alertmanager
  • Experience with log aggregation using Loki, ELK Stack, or equivalent
  • Understanding of distributed systems and microservices architecture
  • Experience with Kubernetes or comparable container orchestration
  • Proficiency in Python, Go, or Bash
  • Strong analytical skills for root-cause analysis and capacity planning
  • Hands-on OpenTelemetry instrumentation experience
  • Exposure to Grafana OnCall, Grafana Incident, or PagerDuty
  • Familiarity with eBPF-based observability tools such as Cilium or Parca
  • Azure or AWS certifications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2009

What We Do

nHRMS is a strategic human-resources partner providing end-to-end solutions for organizations in the United States and India. Its services span executive search and talent acquisition, performance management, HR advisory, organization strategy, HR technology, leadership development, labor-code compliance, workforce productivity, and learning. The firm supports clients across the employee lifecycle, combining people-first consulting with technology-enabled systems to help organizations scale.

Similar Jobs

In-Office
Hyderabad, Telangana, IND
3062 Employees

Micron Technology Logo Micron Technology

Senior /Staff AMS Layout engineer-DPG-LPDDR

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Hyderabad, Telangana, IND
45000 Employees

Micron Technology Logo Micron Technology

Engineer/ Sr Memory Layout Engineer-DPG-LPDDR

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Hyderabad, Telangana, IND
45000 Employees

Micron Technology Logo Micron Technology

Senior Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Hyderabad, Telangana, IND
45000 Employees

Similar Companies Hiring

Empathy Thumbnail
Fintech • Healthtech • HR Tech • Information Technology • Financial Services • Telehealth
IL
200 Employees
Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account