Site Reliability Engineer II - CTJ - Secret

Reposted Yesterday
Be an Early Applicant
Reston, VA, USA
In-Office
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Designs, develops, and operates reliable, scalable software for Microsoft Office 365 government cloud environments. Responsibilities include distributed systems engineering, telemetry analysis, monitoring, automation, deployment, incident response, performance optimization, capacity planning, and root-cause analysis. The engineer participates in on-call rotations, code and design reviews, and collaborates with product engineering and program management teams to improve availability, security, reliability, and operational efficiency.
Summary Generated by Built In
Overview

Do you have a passion for high scale services and working with some of Microsoft’s most critical customers? We’re looking for a Site Reliability Engineer II with the right mix of software development, on-line services experience and passion for quality to envision, design, and deliver Office 365 government cloud service offerings.

Office 365 is at the center of Microsoft’s cloud first, devices first strategy as it brings together cloud versions of our most trusted communication and collaboration products like Exchange, SharePoint, and Teams with our cross-platform desktop suites and mobile apps. The Office 365 Enterprise Cloud team works with Microsoft’s largest enterprise and government customers to deliver features that meet their specific needs and enable cloud adoption. As you would expect, our customers have the highest expectations for feature quality, security, reliability, availability, and performance.

At Microsoft, we can offer you a solid team, exciting challenges, and a fun place to work. The work environment empowers you to have a positive impact on millions of end users. 

The Site Reliability Engineering (SRE) team provides leadership, direction and accountability for application architecture, system design, and end-to-end implementation.

Internal CSP alignment: This role aligns to the Microsoft Career Stage Profile (CSP) for Site Reliability Engineering IC3. The role emphasizes independent execution and service ownership, using telemetry and automation to drive availability, reliability, performance, efficiency, observability, and operational improvements.

As a Site Reliability Engineer II, you will identify and deliver software improvements using your expertise in software development, complexity analysis, and scalable system design. Solid collaboration skills will be required to work closely with other engineering teams to ensure services/systems are highly stable and performant, meeting the expectations of our government customers and users. 

At Microsoft, our mission—to empower every person and every organization on the planet to achieve more—guides how we partner with customers to deliver trusted, impactful solutions. With a growth mindset culture, we innovate responsibly and measure success by shared progress—people, teams, and customers. Join us to do meaningful work that changes the world and helps shape what’s next for everyone.   


Responsibilities

Technical Knowledge and Domain-Specific Expertise

  • Demonstrates expertise in distributed systems design, interactions between cloud technology layers and components, common dependencies at scale, and the code that defines infrastructures. Can identify and recommend configurations optimal of cloud technology solutions and modify the code base that defines systems or cloud technologies to improve the reliability and operability of supported products with minimal guidance from other engineers.
  • Develops an understanding of the code, features, and operations of specific products at scale as required to contribute to incremental improvements in product availability, reliability, efficiency, observability, and/or performance; participates in on-boarding, code/design reviews, and regular meetings with the engineering teams that develop and/or manage those products.
  • Researches and maintains an awareness in industry trends, advances in distributed systems and cloud technologies, new tools, and/or processes for maintaining and improving product availability, reliability, efficiency, observability, and/or performance. Contributes to the implementation of new solutions within their team by identifying ways they can be applied to solve persistent problems.

Contributions to Development and Design

  • Leverages technical expertise in large scale distributed systems and specific products, as well as objective insights drawn from analyses of production telemetry data to suggest changes or add-ons to product features or code to improve the availability, reliability, efficiency, observability, and performance of product components or features supported by their team.
  • Develops and tests basic changes to optimize code and improve the observability, reliability and operability of a defined range of platform, system, or product components or features with direction from other engineers.
  • Engages with product engineering teams by participating code/design reviews, regular meetings, on-call rotations and incident responses throughout product development and operations cycles; leverages technical expertise on underlying systems/platforms and insights drawn from engagements with product engineering teams and telemetry analyses to propose potential improvements in code base and designs across components and features of one or more products.

Driving Operational Excellence

  • Independently develops code or scripts that automate the performance of repetitive and easily scalable operations processes (e.g., monitoring, alerting, deploying products and updates) across components and features of products operating at scale.
  • Leverages technical expertise and telemetry analysis across a range of components and/or features to identify patterns and opportunities to implement configuration and data changes for one or more platforms, systems, or products in production using code, tooling, and automation.
  • Identifies opportunities to leverage existing tools and automation to enable product engineering teams to increase the velocity in which they can reliably and safely implement changes in production; monitors the effects of changes across multiple components or features within a single platform or system.
  • Designs, develops, and maintains telemetry pipelines and monitoring tools that detail operations metrics (e.g., availability, reliability, performance, efficiency) of product components and features operating at scale. Independently performs analyses using existing tools and/or models to identify insights and shares them with product engineering teams to directly contribute to improvements in product development and/or operations; monitors the impact of changes on operations metrics (e.g., Time-to-X).
  • Independently uses existing tools and/or models to troubleshoot problems or flaws affecting the availability, reliability, performance, and/or efficiency of components and features; proposes solutions that will resolve and prevent recurring issues and brings them to the attention of their Site Reliability Engineering (SRE) and/or product engineering teams.
  • Responds to incidents during regular on-call rotations by identifying the level of impact, troubleshooting issues, and deploying appropriate fixes to resolve root cause(s); alerts product teams and owners to major customer impacting issues and escalates resolution of highly impactful issues affecting multiple components or features to other engineers or engineering teams as needed. Shares details related to incidents and their resolution through post-mortem reports and during regular review meetings.
  • Develops alerts and instrumentation across components and features to monitor product capacity and resource demands and analyze telemetry data using existing capacity planning models; draws insights from analyses of capacity and resource data to optimize component and feature code to manage resources and capacity across limited range of use conditions and system parameters.
  • Utilizes insights from performance and resource monitoring tools to identify whether there is a need to optimize the efficiency of component and feature code, or if changes to compute resources are required; models the predicted effect of changes to code and/or compute resources across components or features to document the efficacy of proposed solutions.
  • Shares insights and Preferred practices that can be applied to improve development and operations of system, platform, or product components and features by participating in code/design reviews, incident drills and debriefs, and regular meetings, as well as interactions with more experienced SREs and members of product engineering teams.

Additional Duties

  • Design, develop, and deliver the required software engineering to serve and protect O365 government clouds.
  • Own deployment, availability, reliability, performance and customer escalation targets for sovereign environments.  
  • Proactively identify and reduce issues through design, testing, and implementation of software-based solutions.  
  • Collaborate with Engineering and Program Management partners to translate customer, business, and technical requirements into architectural designs and feature releases.  
  • Drive efficiencies through software improvement and root cause analysis resulting in service delivery, maturity, and scalability.  
  • Work within a highly skilled team of engineers to deliver revolutionary improvements to the cloud and scale them.

Other

  • Embody our culture and values

Qualifications

Required/Minimum Qualifications:

  • Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration
    • OR Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration 
    • OR equivalent experience. 

Other Requirements: 

Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or 
government security screening requirements as required for this role. These requirements include, but 
are not limited to the following specialized security screenings: 

  • The successful candidate must have an active U.S. Government Secret Security Clearance. Ability 
    to meet Microsoft, customer and/or government security screening requirements are required for 
    this role. Failure to maintain or obtain the appropriate clearance and/or customer screening 
    requirements may result in employment action up to and including termination.
  • Clearance Verification: This position requires successful verification of the stated security 
    clearance to meet federal government customer requirements. You will be asked to provide 
    clearance verification information prior to an offer of employment.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud 
    background check upon hire/transfer and every two years thereafter. 
  • Citizenship & Citizenship Verification: This position requires verification of U.S. citizenship due 
    to citizenship-based legal restrictions. Specifically, this position supports United States federal, 
    state, and/or local United States government agency customer and is subject to certain 
    citizenship-based restrictions where required or permitted by applicable law. To meet this legal 
    requirement, citizenship will be verified via a valid passport, or other approved documents, or 
    verified US government Clearance.

This position is part of a team providing operational coverage for Microsoft Government cloud environments. Engineers participate in a scheduled shift rotation designed to provide extended business-hour coverage. Shift assignments may include alternate schedules such as four 10-hour workdays or rotating coverage across multiple daily operating windows.

Preferred Qualifications:

  • Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration
    • OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration
    • OR equivalent experience.
  • 2+ years technical experience working with large-scale cloud or distributed systems.

#DPG #DPGhiring 


Site Reliability Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, or a related field and 2+ years of technical experience in software engineering, network engineering, or systems administration; or a master's degree and 1+ year of experience; or equivalent experience
  • Active U.S. Government Secret security clearance
  • Successful verification of the stated security clearance
  • Pass the Microsoft Cloud Background Check upon hire or transfer and every two years thereafter
  • Verification of U.S. citizenship
  • Master's degree and 3+ years of relevant technical experience, or bachelor's degree and 5+ years of relevant technical experience, or equivalent experience
  • 2+ years of technical experience working with large-scale cloud or distributed systems

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation — Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support — Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support — Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

In-Office or Remote
2 Locations
175633 Employees
133K-251K Annually

TransUnion Logo TransUnion

Consultant

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
3 Locations
13000 Employees
72K-105K Annually

Vantor Logo Vantor

Machine Learning Engineer

Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
In-Office
McLean, VA, USA
2500 Employees
165K-242K Annually
Hybrid
3 Locations
121228 Employees
40K-100K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account