Senior Site Reliability Engineer

Posted 3 Days Ago
Be an Early Applicant
Redmond, WA, USA
In-Office
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Build and improve automation and tooling to detect, analyze, and mitigate live-site incidents for Azure Cosmos DB. Collaborate with engineering and customers to enhance telemetry, observability, and proactive alerting, perform automated root-cause analysis, and influence product architecture to meet strict SLOs.
Summary Generated by Built In
Overview

Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.
Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.


Within Azure Data, the databases team builds and maintains Microsoft's operational Database systems. We store and manage data in a structured way to enable multitude of applications across various industries. We are on a journey to enable developer friendly, mission-critical, AI enabled operational databases across relational, non-relational and OSS offerings.


We believe in making the day in the life of the On-Call Engineer boring while living up to the expectations of a massive cloud service with stringent Service Level Objectives (SLO’s). We do this by thinking differently, stretching ourselves to go all the way to the root of the problem, keeping data in front and center for all our decisions and taking a systems approach for generating outcomes that far exceeds the expectations. Helping attain the aspirational Service Level Objectives (SLO’s) through pragmatic innovation is what sets the SRE’s in Cosmos DB apart. If you share the same purpose, cause and belief and have passion to follow this pursuit, please read the rest of the Job description on what we do, and we would love to have you join us!
Azure Cosmos DB is Microsoft’s next generation of globally distributed, massively scalable, multi-model cloud database service. It is designed to enable developers to build planet-scale applications. Azure Cosmos DB is one of the fastest growing Azure services. Joining the Azure Cosmos DB team is a fantastic opportunity to work with incredibly talented engineers operating like a startup and be at the forefront of building and shaping the Livesite Automation and AI Ops stack in Cosmos DB and lead the path for broader adoption across Microsoft Azure.


Cosmos DB is a database of choice for the spectrum spanning from the hobbyist developer to the largest of Fortune 500 companies. The database provides the data backbone of many critical systems in Health Care, Retail, Telecommunications, IoT and many more where the Service Availability and Latency is paramount. Cosmos DB provides financially backed SLA (service level agreements) around 99.99 Availability and < 10 MS Latency and we are responsible of upholding ourselves to even more stringent Service Level Objectives (SLO) that delight our customers. Other than a resilient and fault tolerant architecture, a key to attaining the SLO’s is automating the root cause analysis and mitigation of issues and a lot of times proactively addressing the issues even before any customer impact. This team supports on building systems where a vast majority of Livesite issues are automatically mitigated without the need for human intervention.
We are looking for a self-driven Senior Site Reliability Engineer (SRE) who likes taking a data driven and systems-based approach to solve Service Reliability problems. You will be responsible for building and optimizing solutions that can analyze massive amounts of telemetry and other Service Health indicators in near real time and perform automated root cause analysis and necessary mitigations to restore SLO’s.


Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Collaborating closely with engineering teams on building and enhancing tooling and automation solutions for faster resolution of issues impacting SLO’s and averting incidents altogether when possible.
  • Collaborating with the customers to understand their pain points around supportability and SLO attainment and formulate strategies for addressing recurring issues in a sustainable way.
  • Communicate on a deeply technical level and be the single point of contact for interfacing with enterprise customers for handling service escalations and driving the issues to resolution.
  • Ability to design and implement any changes to service telemetry for the automation to consume if it is not already available.
  • Enhancing customer facing experience by proactive alerting based on utilization, trends, resource health, etc.
  • Analyze data and provide operational insights into customer experience to design and product teams, so that we can design features with supportability in mind.
  • Embody our culture and values.

Qualifications

Required/Minimum Qualifications:

  • 6+ years technical experience in software engineering, network engineering, or systems administration
    • OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration
    • OR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration.

Other Requirements:


Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:

    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred/Additional Qualifications:


  • 4+ years of experience running large scale cloud services.
  • 2+ years of operational experience in improving Service Reliability, Availability and Performance.
  • Understanding of Observability and MELT implementation patterns for large-scale services.
  • Experience in Logic Apps and authoring Jupyter Notebooks.
  • Experience in analyzing, troubleshooting, and automating root cause analysis and mitigation of incidents impacting large-scale distributed systems.
  • Systematic problem-solving approach, coupled with effective communication skills and a sense of curiosity.
  • Ability to deal with the ambiguity associated with working in a fast-paced environment.
  • Influencing the product architecture and roadmap to make sure the customer-experienced supportability is always a key consideration when evolving the product.

    

 #azdat #azuredata #SRE


Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's degree + 3+ years OR Master's degree + 2+ years
  • Ability to meet Microsoft, customer and/or government security screening requirements
  • Pass Microsoft Cloud background check upon hire/transfer and every two years thereafter
  • 4+ years of experience running large scale cloud services
  • 2+ years of operational experience improving service reliability, availability, and performance
  • Understanding of observability and MELT (metrics, events, logs, traces) implementation patterns for large-scale services
  • Experience analyzing, troubleshooting, and automating root cause analysis and mitigation for incidents in large distributed systems
  • Experience with Logic Apps and authoring Jupyter Notebooks
  • Strong systematic problem-solving, communication skills, and ability to work in ambiguous, fast-paced environments
  • Influence product architecture and roadmap to improve supportability and SLO attainment

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Remote or Hybrid
United States
1750 Employees

Axon Logo Axon

Senior Site Reliability Engineer

Artificial Intelligence • Cloud • Social Impact • Software • Wearables
In-Office
Seattle, WA, USA
2700 Employees
134K-215K Annually

Microsoft Logo Microsoft

Senior Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office
2 Locations
206870 Employees
120K-261K Annually

Microsoft Logo Microsoft

Senior Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office or Remote
2 Locations
206870 Employees
120K-261K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account