We are looking for Incident Management Engineers to support a global engagement focused on incident handling, communication, and post-incident reviews. The role requires strong analytical thinking, composure under pressure, and excellent communication skills to manage critical incidents effectively.
Responsibilities- Own the end-to-end lifecycle of all major P1 and P2 incidents within our 16x7 shift window, ensuring response and resolution milestones strictly adhere to corporate SLAs.
- Root Cause Analysis: Drive rapid technical triage by analyzing telemetry data, system metrics, monitors and logs to isolate the root cause of complex infrastructure and application failures.
- Data-Driven Troubleshooting: Proven ability to quickly interpret telemetry data, consumer lags, and pipeline bottlenecks under high-pressure scenarios to guide engineering teams toward a fix.
- Anomaly Mitigation: Actively monitor for performance anomalies, queue lags, and throughput drops to proactively mitigate downstream service degradation.
- Tool Proficiency: Experience with observability and monitoring platforms such as Datadog, Grafana
Assertive Communication & Stakeholder Alignment
- Maintain clear, concise, and assertive communication under pressure, cutting through technical noise to extract actionable statuses.
- Formulate and broadcast timely business-focused impact statements and progress metrics to executive leadership and client-facing teams.
- Isolate technical internal chat channels from high-level notification streams to keep critical updates data-rich and highly orderly.
Post-Incident Evolution & Continuous Improvement
- Perform rigorous Root Cause Analysis (RCA) once an incident is safely stood down. Facilitate and contribute to blameless Post-Incident Reviews (PIR) to track down systematic process or system vulnerabilities.
- Actively isolate operational bottlenecks and optimize playbooks to continuously improve Mean Time to Mitigate (MTTM) across critical systems
- Maintain structured handoffs between regions (EMEA & APAC)
- Experience in SRE / Incident Management / Production Support
- Strong communication & negotiation skills (must-have)
- Ability to manage high-pressure situations confidently
- Strong problem-solving and analytical mindset
- Technical knowledge on task execution.
- Good eye for details and understanding of workflows
Technical Skills:
- Ability to proactively identify risks using monitoring tools such as DataDog and Grafana dashboards
- Experience in incident response with capability to quickly restore services (restart, patch, or remediate live issues)
- Strong focus on minimizing service downtime across environments
- Hands-on experience supporting both on-premises (Linux environments) and cloud platforms (primarily Azure, with some exposure to GCP)
- Solid understanding of networking concepts and system architecture
- Understanding incident impact and skills to analyze and take decisions based on them.
- Good understanding of Service now, PagerDuty, JIRA, Databricks, Github actions and basics of Docker and Kubernetes.
- Understanding monitoring systems and able to troubleshoot the root cause of issue.
Skills Required
- Experience in SRE, incident management, or production support
- Strong communication and negotiation skills
- Ability to manage high-pressure situations confidently
- Strong problem-solving and analytical mindset
- Technical knowledge of task execution
- Attention to detail and understanding of workflows
- Ability to identify risks using DataDog and Grafana dashboards
- Incident response experience, including restarting, patching, or remediating live issues
- Experience minimizing service downtime across environments
- Hands-on experience supporting Linux environments and cloud platforms, primarily Azure with some GCP exposure
- Understanding of networking concepts and system architecture
- Understanding of ServiceNow, PagerDuty, Jira, Databricks, GitHub Actions, Docker, and Kubernetes
- Understanding of monitoring systems and root cause troubleshooting
What We Do
Choosing a digital partner is about more than capabilities — it’s about collaboration and character. Unrealistic overhauls and off-the-shelf products ignore what matters most — your unique needs, culture, goals, and your legacy data and technology environments. At EXL, our collaboration is built on ongoing listening and learning to adapt our methodologies. We’re your business evolution partner—tailoring solutions that make the most of data to make better business decisions and drive more intelligence into your increasingly digital operations. Whether your goals are scaling the use of AI and digital, redesign operating models, or driving better and faster decisions, we’re here to partner with you to help you gain—and maintain—competitive advantage with efficient, sustainable models at scale. Our expertise in transformation, data science, and change management helps make your business more efficient and effective, improve customer relationships and enhance revenue growth. Instead of focusing on multi-year, resource- and time-intensive platform designs or migrations, we look deeper at your entire value chain to integrate strategies with impact. We use our specialization in analytics, digital interventions, and operations management—alongside deep industry expertise — to deliver solutions that help you outperform the competition. At EXL, it’s all about outcomes—your outcomes—and delivering success on your terms. Share your goals with us and together, we’ll optimize how you leverage data to drive your business forward. For more information, visit www.exlservice.com.







