Senior Site Reliability Engineer - CTJ - Poly

Posted Yesterday
Be an Early Applicant
3 Locations
In-Office or Remote
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Owns reliability and operations for Azure Resource Manager and large-scale distributed cloud services. Responsibilities include on-call incident response, monitoring, observability, automation, performance optimization, service lifecycle ownership, coding, design documentation, and cross-functional delivery. The role requires improving availability, security, efficiency, and customer support across cloud environments while guiding engineers and maintaining service parity with the commercial cloud.
Summary Generated by Built In
Overview

Want to work on the control plane for Azure itself? Our team owns Azure Resource Manager, the system every customer and service depends on to deploy, manage, and organize resources. We build large-scale, fault-tolerant distributed systems that define how Azure operates, from resource lifecycle and access control to global inventory and governance.

Our team fosters a collaborative environment and builds upon each other’s ideas, to deliver world-class customer value at a rapid pace. We empower engineers to deliver creative solutions through bottoms-up innovation. This is a fun environment and a great opportunity to work on something highly strategic to Microsoft and extremely relevant in the industry. We’re looking for a Senior Site Reliability Engineer and a leader passionate about delivering value to customers in mission critical environments, who enjoys a growth hacking culture, and is eager to play a part in one of the most important long games for Microsoft.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of psychological safety where everyone can thrive at work.


Responsibilities
  • Acts as a Designated Responsible Individual (DRI) and guides other engineers by developing and following the playbook, working on call to monitor system/product/service for degradation, downtime, or interruptions, alerting stakeholders about status and initiates actions to restore system/product/service for simple and complex problems when appropriate.
  • Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of service fabric services while also driving consistency in monitoring and operations at scale
  • Drives development of design documents for a product, application, service, or platform.
  • Creates, implements, optimizes, debugs, refactors, and reuses code to establish and improve performance and maintainability, effectiveness, and return on investment (ROI).
  • Leverages subject-matter expertise of product features and partners with appropriate stakeholders (e.g., project managers) to drive a workgroup's project plans, release plans, and work items.
  • Take full ownership of assigned services, actively contributing to its enhancement across all cloud environments. Ensure the service maintains parity with the commercial cloud, delivering high support standards for customers. Participate in the service lifecycle, including design, development, deployment, and maintenance. Collaborate with cross-functional teams to uphold the highest standards of quality and performance. Engage in continuous improvement initiatives to enhance the service's capabilities and user experience.
  • Identify opportunities for automation and optimization within the cloud to better support customers. This includes evaluating current processes and workflows to pinpoint inefficiencies and areas for enhancement. Design and implement automation solutions to streamline operations, reduce manual effort, and improve overall service delivery. Focus on optimizing existing systems and processes to boost performance and customer satisfaction.
  • Embody our culture and values.

  •  

Qualifications
Required Qualifications:
  • Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 

Other Requirements: 

Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:  

  • The successful candidate must have an active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI) with Polygraph. Ability to meet Microsoft, customer and/or government security screening requirements are required pre-offer and post-hire for this role. Failure to maintain or obtain the appropriate U.S. Government clearance and/or customer screening requirements may result in employment action up to and including termination. 
  • Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.   
  • Citizenship & Citizenship Verification: This position requires verification of U.S. citizenship due to citizenship-based legal restrictions. Specifically, this position supports United States federal, state, and/or local United States government agency customer and is subject to certain citizenship-based restrictions where required or permitted by applicable law. To meet this legal requirement, citizenship will be verified via a valid passport, or other approved documents, or verified US government Clearance 

Preferred Qualifications:

  • Doctorate Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
  • 3+ years technical experience working with large-scale cloud or distributed systems.
  • Experience writing scripts and functional programming code to automate tasks, using languages such as Python, JavaScript, or Shell scripting.
  • Experience developing end-to-end technical expertise in the architecture, code, features, and operations of specific products as required to implement improvements in product availability, security, quality, observability, reliability, efficiency, observability, and/or performance.
  • Experience driving code/design reviews with the engineering teams that develop and/or manage those products and shares learnings and recommendations across engineering teams working on related products within their organization and other organizations as relevant. 
  • Knowledge of distributed systems.
  • Highly effective written and oral communication skills.


Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Master’s degree in Computer Science, Information Technology, or a related field plus 2+ years of technical experience in software engineering, network engineering, or systems administration; or bachelor’s degree in a related field plus 4+ years of technical experience; or equivalent experience
  • Active U.S. Government Top Secret clearance with access to Sensitive Compartmented Information based on a Single Scope Background Investigation with Polygraph
  • Ability to meet Microsoft, customer, and government security screening requirements
  • Successful verification of the stated security clearance
  • Pass the Microsoft Cloud background check upon hire or transfer and every two years thereafter
  • U.S. citizenship verification
  • Doctorate degree plus 3+ years, master’s degree plus 6+ years, or bachelor’s degree plus 8+ years of relevant technical experience; or equivalent experience
  • 3+ years of experience working with large-scale cloud or distributed systems
  • Experience writing scripts and functional programming code using Python, JavaScript, or Shell scripting
  • End-to-end technical expertise in product architecture, code, features, and operations
  • Experience driving code and design reviews with engineering teams
  • Knowledge of distributed systems
  • Highly effective written and oral communication skills

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Pfizer Logo Pfizer

HSS-Health and Science Specialist - Southern Maryland, MD

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
3 Locations
121990 Employees
94K-165K Annually

Optum Logo Optum

RN Case Manager - HouseCalls - Remote - EST or CST - Compact license Required

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Columbia, MD, USA
160000 Employees
60K-107K Annually

DraftKings Logo DraftKings

Associate, Regulation

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Remote or Hybrid
United States
6400 Employees
56K-70K Annually

MetLife Logo MetLife

Production Support Analyst - CRM & Workflow

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
85K-110K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account