Senior Site Reliability Engineer - CTJ - Poly

Posted Yesterday
Be an Early Applicant
3 Locations
In-Office or Remote
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Lead reliability and availability for large-scale Azure services: act as DRI, run on-call, design and implement improvements, automate operations, develop design docs and code, collaborate across teams, and drive service lifecycle and incident restoration for cloud-hosted services.
Summary Generated by Built In
Overview

Do you want to be at the heart of cloud computing? The Compute team is at the core of Azure and is growing incredibly fast. We build and manage fault tolerant distributed systems on top of commodity datacenter hardware, to deliver an infrastructure for hosting customer applications. The platform is at the core of Azure that provides millions of virtual machines for customers to run their workload in the cloud.

Our team fosters a collaborative environment and builds upon each other’s ideas, to deliver world-class customer value at a rapid pace. We empower engineers to deliver creative solutions through bottoms-up innovation. This is a fun environment and a great opportunity to work on something highly strategic to Microsoft and extremely relevant in the industry. We’re looking for a Senior Site Reliability Engineer and a leader passionate about delivering value to customers in mission critical environments, who enjoys a growth hacking culture, and is eager to play a part in one of the most important long games for Microsoft.


Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of psychological safety where everyone can thrive at work.


Responsibilities
  • Acts as a Designated Responsible Individual (DRI) and guides other engineers by developing and following the playbook, working on call to monitor system/product/service for degradation, downtime, or interruptions, alerting stakeholders about status and initiates actions to restore system/product/service for simple and complex problems when appropriate.
  • Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of service fabric services while also driving consistency in monitoring and operations at scale
  • Drives development of design documents for a product, application, service, or platform.
  • Creates, implements, optimizes, debugs, refactors, and reuses code to establish and improve performance and maintainability, effectiveness, and return on investment (ROI).
  • Leverages subject-matter expertise of product features and partners with appropriate stakeholders (e.g., project managers) to drive a workgroup's project plans, release plans, and work items.
  • Take full ownership of assigned services, actively contributing to its enhancement across all cloud environments. Ensure the service maintains parity with the commercial cloud, delivering high support standards for customers. Participate in the service lifecycle, including design, development, deployment, and maintenance. Collaborate with cross-functional teams to uphold the highest standards of quality and performance. Engage in continuous improvement initiatives to enhance the service's capabilities and user experience.
  • Identify opportunities for automation and optimization within the cloud to better support customers. This includes evaluating current processes and workflows to pinpoint inefficiencies and areas for enhancement. Design and implement automation solutions to streamline operations, reduce manual effort, and improve overall service delivery. Focus on optimizing existing systems and processes to boost performance and customer satisfaction.
  • Embody our culture and values.


Qualifications
Required Qualifications:
  • Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 

Other Requirements: 

Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:  

  • The successful candidate must have an active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI) with Polygraph. Ability to meet Microsoft, customer and/or government security screening requirements are required pre-offer and post-hire for this role. Failure to maintain or obtain the appropriate U.S. Government clearance and/or customer screening requirements may result in employment action up to and including termination. 
  • Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.   
  • Citizenship & Citizenship Verification: This position requires verification of U.S. citizenship due to citizenship-based legal restrictions. Specifically, this position supports United States federal, state, and/or local United States government agency customer and is subject to certain citizenship-based restrictions where required or permitted by applicable law. To meet this legal requirement, citizenship will be verified via a valid passport, or other approved documents, or verified US government Clearance 

Preferred Qualifications:

  • Doctorate Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
     
  • 3+ years technical experience working with large-scale cloud or distributed systems.
  • Experience writing scripts and functional programming code to automate tasks, using languages such as Python, JavaScript, or Shell scripting.
  • Experience developing end-to-end technical expertise in the architecture, code, features, and operations of specific products as required to implement improvements in product availability, security, quality, observability, reliability, efficiency, observability, and/or performance.
  • Experience driving code/design reviews with the engineering teams that develop and/or manage those products and shares learnings and recommendations across engineering teams working on related products within their organization and other organizations as relevant. 
  • Knowledge of distributed systems.
  • Highly effective written and oral communication skills.

Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
  • Active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI) with Polygraph.
  • U.S. citizenship and ability to verify citizenship via passport or approved documents.
  • Ability to meet Microsoft, customer and/or government security screening requirements pre-offer and post-hire, including Microsoft Cloud background check (and recheck every two years).
  • Experience working with large-scale cloud or distributed systems (3+ years).
  • Experience writing scripts and functional programming code to automate tasks using Python, JavaScript, or Shell scripting.
  • Experience developing end-to-end technical expertise in product architecture, code, features, and operations to improve availability, security, observability, reliability, efficiency, and performance.
  • Experience driving code and design reviews and sharing recommendations across engineering teams.
  • Knowledge of distributed systems.
  • Highly effective written and oral communication skills.

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Toast Logo Toast

Principal Software Engineer

Cloud • Fintech • Food • Information Technology • Software • Hospitality
Remote
US
5000 Employees
283K-453K Annually

General Motors Logo General Motors

Account Executive

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
United States
165000 Employees
106K-141K Annually

Collectors Logo Collectors

Senior Platform Engineer

Consumer Web • eCommerce • Machine Learning • Software • Sports • Analytics
In-Office or Remote
2 Locations
2246 Employees
140K-200K Annually

DFIN Logo DFIN

Account Executive

Fintech • Software
Remote or Hybrid
United States
1750 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account