Site Reliability Engineer - CTJ - Poly

Reposted 3 Days Ago
Be an Early Applicant
2 Locations
In-Office
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Lead and develop SRE practices for secure, highly regulated Microsoft CISO services. Architect and automate hybrid/cloud infrastructure, manage petabyte-scale data platforms and pipelines, implement IaC and disaster recovery, deliver telemetry and automation, participate in on-call rotations, and mentor engineers to improve reliability, diagnosability, security, and compliance.
Summary Generated by Built In
Overview

We are seeking a Senior Site Reliability Engineer to lead a team that builds and operates Microsoft CISO security engineering services in highly regulated environments, including U.S. Government Cloud deployments. In this space, success requires both operational rigor and strong software engineering fundamentals, maintainable code, extendable design, robust telemetry, and disciplined lifecycle practices that make reliability a built-in feature. 
 
This role is rooted in software engineering as a reliability lever. You will work with teams that deliver production code, automation, and self-healing capabilities, and partner with feature engineering teams to bake in reliability, diagnosability, security, and compliance from design through operations. You will help operate and evolve large-scale enterprise applications, and multi-petabyte data platforms where availability, resilience, and uptime are mission critical. You will amplify impact by developing engineers, setting up reliability strategies, and influencing how services are built and run across organizational boundaries. 


Responsibilities

Responsibilities: 

  • Write secure, high-quality code that is maintainable, scalable, and performant. 

  • Architect, implement, and optimize hybrid and cloud infrastructure using Infrastructure as Code (e.g., Containers, Bicep, Terraform, AKS etc.) to improve availability, scale, security, and operational efficiency. 

  • Design and implement data governance, storage, backup, and disaster recovery for a multi-petabyte Azure environment, ensuring integrity, security, and performance. 

  • Build and operate large-scale data pipelines and data transformations to support analytics, governance, and operational needs. 

  • Evaluate emerging engineering tools and practices and incorporate them into the roadmap to continuously improve efficiency, reliability, and scale. 

  • Deliver automation to improve service health, manageability, reliability, telemetry, and alerting, with a focus on resiliency. 

  • Create and maintain clear technical documentation and design specifications aligned with best practices. 

  • Partner with engineering, project management, and operations to evolve services and optimize infrastructure in support of organizational goals. 

  • Participate in an on-call rotation to operate live services; troubleshoot and mitigate complex issues, escalate as needed, and write post-incident reviews to share learnings. 

  • Identify opportunities for automation using scripts, pipelines, policydriven guardrails, or AIenabled tooling to reduce manual toil and increase engineering productivity. 


Qualifications

Required/minimum qualifications:

Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
 
Other requirements:
Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:  
  • The successful candidate must have an active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI) with Polygraph. Ability to meet Microsoft, customer and/or government security screening requirements are required pre-offer and post-hire for this role. Failure to maintain or obtain the appropriate U.S. Government clearance and/or customer screening requirements may result in employment action up to and including termination. 
  • Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
  • Citizenship & Citizenship Verification: This position requires verification of U.S. citizenship due to citizenship-based legal restrictions. Specifically, this position supports United States federal, state, and/or local United States government agency customer and is subject to certain citizenship-based restrictions where required or permitted by applicable law. To meet this legal requirement, citizenship will be verified via a valid passport, or other approved documents, or verified US government Clearance.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.   
Additional or preferred qualifications:
Doctorate Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
  • 4+ years of experience building, deploying, and operating containerized applications and infrastructure as code (e.g., Docker, Kubernetes, Azure Container Apps/AKS/ACI, Terraform, Azure Bicep, ARM templates). 

  • 4+ years of experience writing and maintaining scripts for deployment, orchestration, and automation (e.g., PowerShell, Python, Bash). 

  • Experience working with large datasets, data pipelines, and data transformation patterns (batch and/or streaming). 

  • Experience with one or more major cloud platforms (Azure, AWS, or Google Cloud). 
  • Hands-on experience with Azure services and infrastructure (e.g., ARM templates, IaaS, VMs, Key Vault, Event Hubs, Synapse, Spark/Hadoop), or equivalent services in AWS or Google Cloud. 

  • Familiarity with data pipeline and transformation tooling (e.g., Spark, Hadoop) and operating at scale. 

  • Familiarity with large-scale Microsoft enterprise services (e.g., Microsoft 365: Exchange, SharePoint, Skype, Teams). 

  • Familiarity with petabyte-scale datasets and building reliable data pipelines and transformations that support mission-critical services. 

  • Proficiency in at least one programming language (e.g., C# or Java) and scripting languages such as PowerShell, Bash, and Python. 


Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Degree and experience: Masters in CS/IT + 2+ years OR Bachelors in CS/IT + 4+ years, or equivalent experience.
  • Active U.S. Government Top Secret Clearance with SCI based on SSBI and Polygraph.
  • U.S. citizenship verification (must be U.S. citizen).
  • Ability to meet Microsoft, customer, and government security screening requirements pre-offer and post-hire.
  • Pass Microsoft Cloud background check upon hire and every two years thereafter.
  • Participate in on-call rotation and perform incident troubleshooting, mitigation, and post-incident reviews.
  • 4+ years building, deploying, and operating containerized applications and IaC (Docker, Kubernetes, AKS, ACI, Terraform, Azure Bicep, ARM templates).
  • 4+ years writing and maintaining deployment/orchestration/automation scripts (PowerShell, Python, Bash).
  • Experience with large datasets, data pipelines, transformations, and operating at scale (Spark, Hadoop, Synapse).
  • Experience with major cloud platforms (Azure, AWS, or Google Cloud) and hands-on Azure services (VMs, Key Vault, Event Hubs, Synapse, Spark/Hadoop).
  • Proficiency in at least one programming language (C# or Java) and scripting languages (PowerShell, Bash, Python).
  • Familiarity with large-scale Microsoft enterprise services (Microsoft 365: Exchange, SharePoint, Skype, Teams).
  • Experience designing data governance, backup, and disaster recovery for multi-petabyte environments.
  • Doctorate or advanced degree in CS/IT and additional years of technical experience (preferred senior experience levels).

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Microsoft Logo Microsoft

Senior Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office or Remote
3 Locations
206870 Employees
120K-261K Annually

Microsoft Logo Microsoft

Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office or Remote
4 Locations
206870 Employees
102K-219K Annually

Microsoft Logo Microsoft

Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office or Remote
3 Locations
206870 Employees
120K-261K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account