Senior Observability Engineer / Platform Engineer

Posted 19 Days Ago
Be an Early Applicant
Hiring Remotely in Chennai, Tamilnadu, IND
In-Office or Remote
Senior level
Information Technology • Business Intelligence • Consulting
The Role
Designs and maintains enterprise observability platforms for metrics, logs, traces, and events across Kubernetes and public cloud environments. Builds dashboards, SLOs, SLIs, alerts, and service health frameworks; supports incident response, troubleshooting, and root cause analysis. Automates platform onboarding through infrastructure-as-code and CI/CD, while partnering with engineering, SRE, platform, and operations teams to improve reliability. Also contributes to distributed tracing, AIOps, intelligent alerting, and automated remediation initiatives.
Summary Generated by Built In

This is a remote position.

Overview :

We are looking for a hands-on Observability Engineer with strong experience in cloud-native platforms, monitoring, logging, alerting, and operational excellence. The ideal candidate will have experience building and managing enterprise observability solutions across Kubernetes and public cloud environments, enabling proactive monitoring, incident response, and reliability engineering practices.

Key Responsibilities

Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events. Build observability solutions using tools such as Prometheus, Grafana, OpenSearch, Splunk, Elastic, Datadog, Dynatrace, New Relic, or equivalent platforms. Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks. Integrate monitoring and observability capabilities within Kubernetes and cloud-native environments. Enable incident management, root cause analysis, and operational troubleshooting through observability best practices. Automate monitoring configuration and platform onboarding using Infrastructure-as-Code and CI/CD pipelines. Collaborate with engineering, platform, SRE, and operations teams to improve reliability, performance, and availability. Support observability maturity initiatives including distributed tracing, AIOps, and intelligent alerting.



Requirements

Required Skills

  • 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, or Observability Engineering.
  • Strong expertise in: Prometheus Grafana Splunk / OpenSearch / Elastic Distributed tracing solutions (Jaeger, Tempo, OpenTelemetry, etc.) Good understanding of Kubernetes, Docker, and containerized workloads.
  • Experience with AWS, Azure, or GCP environments. Knowledge of incident management, alert tuning, and troubleshooting production environments. Experience with scripting and automation using Python, Shell, or similar languages.
  • Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or ArgoCD. Preferred Skills Exposure to SRE practices, error budgets, SLO/SLI frameworks.
  • Experience with AIOps, automated remediation, or intelligent incident response. Knowledge of OpenTelemetry implementation.
  • Working experience in enterprise-scale production platforms. Exposure to security and governance considerations in cloud-native environments.


Benefits

Diversity Inclusion:

At Exavalu, we are committed to building a diverse and inclusive workforce. We welcome applications for employment from all qualified candidates, regardless of race, color, gender, national or ethnic origin, age, disability, religion, sexual orientation, gender identity or any other status protected by applicable law. We nurture a culture that embraces all individuals and promotes diverse perspectives, where you can make an impact and grow your career.

Exavalu also promotes flexibility  depending on the needs of employees, customers and the business. It might be part-time work, working outside normal 9-5 business hours or working remotely. We also have a welcome back program to help people get back to the mainstream after a long break due to health or family reasons.



Skills Required

  • 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, or Observability Engineering
  • Strong expertise with Prometheus, Grafana, and Splunk, OpenSearch, or Elastic
  • Experience with distributed tracing solutions such as Jaeger, Tempo, or OpenTelemetry
  • Understanding of Kubernetes, Docker, and containerized workloads
  • Experience with AWS, Azure, or GCP environments
  • Knowledge of incident management, alert tuning, and production troubleshooting
  • Experience with scripting and automation using Python, Shell, or similar languages
  • Experience with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or ArgoCD
  • Exposure to SRE practices, error budgets, and SLO/SLI frameworks
  • Experience with AIOps, automated remediation, or intelligent incident response
  • Knowledge of OpenTelemetry implementation
  • Experience supporting enterprise-scale production platforms
  • Exposure to security and governance considerations in cloud-native environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Chennai, Tamilnadu
295 Employees
Year Founded: 2018

What We Do

Exavalu leads the way in business and technology consulting and solution delivery, focusing on specialized digital transformation across the Insurance, Healthcare & Life Sciences, and Nonprofit industries. As an award-winning digital transformation advisor, we take pride in delivering leading-edge solutions to clients worldwide. Led by industry veterans and technology experts, we bring our deep understanding in customer experience, process automation, digital engineering, cloud, data and AI with a very strong industry core foundation. Drawing on our industry insights and technological acumen, we enable clients to drives growth, profitability, and cost optimization. We mitigate risks and elevate customer experience through our comprehensive suite of services, including strategic advisory, expert-led technology [build and maintenance], and best-in-class products and solutions. Our unwavering focus on understanding the industry's hard problems and finding the right solution has garnered recognition as the preferred ally for organizations seeking transformative growth.

Similar Jobs

Capco Logo Capco

IT Control Testing

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Atlassian Logo Atlassian

Program Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
India
11000 Employees

The Aerospace Corporation Logo The Aerospace Corporation

Integrity &Compl Specialist II

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
India
4600 Employees

Magnite Logo Magnite

Security Engineer

AdTech • Big Data • Digital Media • Software
Remote or Hybrid
India
950 Employees

Similar Companies Hiring

Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account