Site Reliability Engineer 2

Reposted 2 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Provide on-call incident triage and first-line response for data integration services. Build AI-driven agents and automation for triage, routing, known-issue matching, incident summarization, and postmortem generation. Improve observability, metrics (FTM, deflection, time-to-triage), and troubleshoot via automated tooling while enforcing operational standards.
Summary Generated by Built In
Overview

Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.
Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.

Within Microsoft Fabric, the Data Integration team enables organizations to connect, move, and shape data across their entire data estate. As data is generated across applications, devices, and systems, our integration capabilities make it easy to ingest, transform, and prepare data for downstream use. Fabric’s unified integration experience simplifies how customers bring data together, ensuring it is trusted, usable, and ready to power analytics, real-time insights, and AI scenarios.

The Customer Data Integration (CDI) organization within Power Query is building a new Live site engineering team that serves as the first line of defense for one of Microsoft's most critical data integration services. You will own incident triage and response across Power BI, Fabric, Power Query Online, Gateway, and hundreds of data connectors,  and you will build the agentic automation that makes that operation increasingly self-driving.
This is not a passive monitoring role. You will be on call, triaging real incidents, and ensuring uptime for business-critical services. You will simultaneously build the intelligent agents, automation pipelines, and tooling that reduce manual effort with every iteration. Your goal is to make yourself more effective over time by engineering your way out of repetitive work.

We do not just value differences or different perspectives. We seek them out and invite them in so we can tap into the collective power of everyone in the company. As a result, our customers are better served.


Responsibilities
  • Incident triage and first-line response: Provide on-call coverage for incoming incidents across CDI services. Perform initial investigation, severity assessment, and routing to owning engineering teams.
  • Agentic triage system development: Build and extend AI-driven agents that ingest ICM alerts, correlate with recent deployments and feature flag rollouts, check known-issue databases, and produce initial assessments with suggested severity and owning team.
  • TSG and known-issue matching: Develop automation that matches incoming incidents to relevant Troubleshooting Guides (TSGs) and known issues across Fabric and Power Platform — reducing investigation time and enabling faster resolution.

Auto-routing and classification: Configure and extend ICM routing rules and build intelligent classification systems based on service tree, alert signatures, and historical patterns.

  • Incident lifecycle automation: Build agents for incident summarization, customer communications drafting, postmortem generation, and reporting, replacing manual authoring with AI-assisted workflows requiring human judgment only for high-severity incidents.
  •  Metrics and continuous improvement: Measure and improve first-time mitigation (FTM) rates, incident deflection rates, and time-to-triage. Use data to identify patterns, propose new TSGs, and drive systemic improvements.
  • CSS quality bar: Defend the standard for incoming escalations by enforcing organizational and process standards
  • Embody our culture and values

Qualifications

Required/Minimum Qualifications

  • 4+ years of software engineering experience in site reliability, Live site operations, or incident management for cloud services.
  • Good programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto.
  •  Experience with incident management systems and workflows (ICM, PagerDuty, ServiceNow, or similar).
  •  Experience with monitoring, alerting, and observability systems (Kusto, Geneva, Grafana, or similar).
  • Ability to work in an on-call rotation across time zones in a geographically distributed team.
  • Experience interface with engineers, leadership, support, and customers.

Job Requirements: Other & Additional

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check:
  • This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred/Additional Qualifications

  • Experience building AI/ML-driven automation, agents, or intelligent workflows (e.g., using LLMs, Copilot extensibility, MCP servers, or agentic frameworks).
  • Familiarity with Live site ecosystem management (including log traversal, incident management, telemetry analysis, etc.)
  •  Experience with Azure, Power BI, and Fabric services.
  •  Experience with Troubleshooting Guide (TSG) authoring and incident pattern analysis.
  •  Understanding of SLA management, customer communications, and escalation workflows for cloud services.

Benefits/perks listed below may vary depending on the nature of your employment with Microsoft and the country where you work.

#azdat

#azuredata

#dataintegration #powerbi #powerquery #dataflows #fabric


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • 4+ years of software engineering experience in site reliability, live site operations, or incident management for cloud services.
  • Programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto.
  • Experience with incident management systems and workflows (ICM, PagerDuty, ServiceNow, or similar).
  • Experience with monitoring, alerting, and observability systems (Kusto, Geneva, Grafana, or similar).
  • Ability to work in an on-call rotation across time zones in a geographically distributed team.
  • Experience interfacing with engineers, leadership, support, and customers.
  • Ability to meet Microsoft Cloud background check and related security screening requirements.
  • Experience building AI/ML-driven automation, agents, or intelligent workflows (e.g., using LLMs, Copilot extensibility, MCP servers, or agentic frameworks).
  • Familiarity with live site ecosystem management, log traversal, telemetry analysis, and incident pattern analysis.
  • Experience with Azure, Power BI, and Fabric services.
  • Experience with Troubleshooting Guide (TSG) authoring and SLA management, customer communications, and escalation workflows.

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Zeta Global Logo Zeta Global

Analyst - SEC Reporting

AdTech • Artificial Intelligence • Marketing Tech • Software • Analytics
Easy Apply
Hybrid
Hyderabad, Telangana, IND
2429 Employees

Capco Logo Capco

Managing Principal - Data

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Crunchyroll Logo Crunchyroll

Senior Manager, Data Science and Machine Learning

Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Hybrid
Hyderabad, Telangana, IND
1300 Employees

Optum Logo Optum

Director - Back Office RCM Operations

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account