Senior Software Engineer

Posted 3 Days Ago
Be an Early Applicant
Redmond, WA, USA
In-Office
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Design and lead architecture for agentic reliability and observability platforms. Build automated detection, triage, mitigation, and safe remediation systems. Drive operational standards (SLOs, alert quality, incident automation), influence partner teams, run live-site investigations, convert reliability gaps into platform investments, and mentor engineers on monitoring, incident response, and AI-assisted automation.
Summary Generated by Built In
Overview

Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.

Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.

Within Microsoft Fabric, the Azure Monitor team builds services that enable customers to monitor, detect, troubleshoot, and mitigate issues with their services through an increasingly agentic experience. Azure Monitor includes Log Analytics, Application Insights, Container Insights, Hosted Prometheus, Azure Managed Grafana, and more. Additionally, Azure Monitor is the platform upon which Microsoft Sentinel is built. We have a multi-billion dollar business that is growing rapidly, and we run some of the world’s highest scale observability services both for Microsoft internally and for our external customers, processing over 1.5 Exabytes of logs daily and tracking over 100 billion active metrics.​

The Azure Monitor Site Reliability Team is hiring a Senior Software Engineer to drive architecture and delivery of agentic reliability platform capabilities for Microsoft Monitoring solutions. This role defines and builds platform systems for monitoring intelligence, telemetry, diagnostics, safe remediation, and operational automation used by external customers and internal Microsoft engineering teams. 

​​We do not just value differences or different perspectives. We seek them out and invite them in so we can tap into the collective power of everyone in the company. As a result, our customers are better served.


Responsibilities
  • Define architecture for agentic reliability systems spanning monitoring, telemetry, incident management, service topology, deployment signals, and operational knowledge.
  • Lead platform capabilities for automated detection, triage, root-cause assistance, mitigation recommendations, safe execution, and post-incident learning.
  • Establish engineering standards for safe agentic operations, including identity, access, compliance, rollback, auditability, change management, and human escalation.
  • Influence service teams to adopt consistent monitoring, SLOs, alert quality, incident automation, live-site readiness, and operational excellence practices.
  • Identify high-impact reliability gaps and convert them into platform investments, architectural improvements, and reusable automation.
  • Lead complex live-site investigations and drive systemic reliability improvements from incident patterns and customer-impact data.
  • Mentor engineers and shape long-term technical direction across monitoring, observability, incident response, and AI-assisted and agentic automation. 
  • Embody our culture and values 

Qualifications
Required Qualifications
Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
 

Other Requirements: 

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check:

  • This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred Qualifications

  • Experience building production-scale platforms, cloud services, distributed systems, or reliability automation.
  • Experience architecting complex systems across service boundaries and driving execution across partner teams without direct authority.
  • Experience with observability architecture, monitoring systems, incident response, service health modeling, operational automation, and production debugging.
  • Solid judgment around production safety, automation risk, customer impact, security, privacy, compliance, and responsible AI-assisted and agentic automation.
  • Experience leading agentic automation, AI-assisted diagnostics, autonomous remediation, intelligent operations, or reliability platform efforts.Experience with Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Azure Resource Graph, Azure DevOps, GitHub, or similar monitoring and observability ecosystems.
  • Experience creating organization-level reliability metrics such as SLO compliance, alert quality, time to detect, time to mitigate, human effort saved, automation coverage, and incident recurrence.
  • Experience mentoring senior engineers and defining technical strategy across monitoring, observability, incident response, and production engineering disciplines.




#azdat 

#azuredata 

​#Azure #AzureMonitor #Monitor #LogAnalytics #Observability #SRE #SiteReliabilityEngineering​ 


Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor's Degree in Computer Science or related technical field and 4+ years technical engineering experience coding in languages such as C, C++, C#, Java, JavaScript, or Python (or equivalent experience).
  • Ability to meet Microsoft, customer and/or government security screening requirements and pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
  • Experience building production-scale platforms, cloud services, distributed systems, or reliability automation.
  • Experience architecting complex systems across service boundaries and driving execution across partner teams without direct authority.
  • Experience with observability architecture, monitoring systems, incident response, service health modeling, operational automation, and production debugging.
  • Experience with AI-assisted diagnostics, autonomous remediation, intelligent operations, or reliability platform efforts.
  • Experience with Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Azure Resource Graph, Azure DevOps, or GitHub (or similar monitoring/observability ecosystems).
  • Experience creating organization-level reliability metrics (SLO compliance, alert quality, time to detect/mitigate, automation coverage, incident recurrence).
  • Experience mentoring senior engineers and defining technical strategy across monitoring, observability, incident response, and production engineering disciplines.

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

CrowdStrike Logo CrowdStrike

Senior Software Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
5 Locations
11000 Employees
140K-215K Annually

Metropolis Technologies Logo Metropolis Technologies

Senior Software Engineer

Artificial Intelligence • Computer Vision • Machine Learning • Payments • Real Estate • PropTech
Easy Apply
In-Office
Seattle, WA, USA
23100 Employees
170K-200K Annually

Motive Logo Motive

Senior Software Engineer

Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Easy Apply
In-Office
2 Locations
4000 Employees
140K-193K Annually

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
155K-208K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account