- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and manage error budgets across services.
- Lead incident management for critical production issues – drive root cause analysis (RCA) and postmortems.
- Create and maintain runbooks and standard operating procedures for high availability services.
- Design and implement observability frameworks using ELK, Prometheus, and Grafana; drive telemetry adoption.
- Coordinate cross-functional war-room sessions during major incidents and maintain response logs.
- Develop and improve automated system recovery, alert suppression, and escalation logic.
- Use GCP tools like GKE, Cloud Monitoring, and Cloud Armor to improve performance and security posture.
- Collaborate with DevOps and Infrastructure teams to build highly available and scalable systems.
- Analyze performance metrics and conduct regular reliability reviews with engineering leads.
- Participate in capacity planning, failover testing, and resilience architecture reviews.
- SRE, DevOs, Infrastructure, and Engineering Teams
- Engineering Leads and Technical Architects
- Cloud and Platform Engineering Teams
- Cross-functional teams during critical production incidents
- Security and Operations teams
- Product and Technology stakeholders
- Teams focused on observability, automation, and platform resilience
- 6–8 years of experience in SRE or Infrastructure Engineering.
- Strong hands-on experience with GCP, GKE, Kubernetes, Docker, Terraform, and Helm.
- Proficiency in Prometheus, Grafana, ELK, and Datadog for observability and monitoring.
- Strong understanding of SLIs, SLOs, error budgets, incident management, RCA, and on-call operations.
- Experience in designing and maintaining highly available, scalable, and resilient cloud-native systems.
- Strong problem-solving, troubleshooting, and communication skills.
- Experience with PagerDuty/OpsGenie and automated incident response.
- Familiarity with programming languages
- Define product strategy by connecting the dots
- Make product decisions without ambiguity in collaboration with the team
- Manage important stakeholders across the organization while leading initiatives
- Decision-making and accountability/ownership is critical
- One point of executive customer engagement for a large portfolio
- Contribute effectively outside your comfort zone
- Opportunity to work on global projects and Fortune 500 clients
- Exposure to cutting-edge technologies
- Strong learning, mentorship, and career growth programs
- Collaborative and innovation-driven work culture
Skills Required
- 6-8 years of SRE or infrastructure engineering experience in cloud-native environments
- GCP (GKE, Load Balancing, VPN, IAM)
- Prometheus
- Grafana
- ELK
- Datadog
- Kubernetes
- Docker
- On-call incident management, RCA, SLIs/SLOs
- Terraform
- Helm
- PagerDuty
- OpsGenie
- GCP Monitoring
- Skywalking
- Service Mesh
- API Gateway
- GCP Spanner
- MongoDB (basic)
- Cloud Armor
What We Do
@TechBlocks we power the software defined industries (SDI) of today and tomorrow. We are a software engineering and consulting firm. We build modern digital value chains and businesses reimagined to create frictionless experiences for innovative monetization methods and drive unforeseen efficiencies. We are known to build world class custom platforms and products that are cloud native for some of the worlds largest brands. We are the go to technology partners for born in digital businesses that grew with us from "Concept to Commercialization" and have revenues between $100M - $10B. We help modern businesses transition just from a technology outsourcing mentality to help create globally distributed digital COEs and mature them. Our converged COEs that we create in partnership with our clients help power software factories that are extremely dynamic. We have created modern digital COEs and factories that are created with a single minded goal to future proof our clients businesses. Everything we do is centred around two philosophies and practices - Design Thinking and Lean Engineering. Whether it is building digital commerce platforms, marketplace for worlds largest retailers or smart utilities applications and products or digital health products/platforms that power wearables, patches or devices across healthcare landscape; we do it all with speed and sophistication that is unmatched in the industry









