Responsibilities
- Live Site Operations: Serve as a Designated Responsible Individual (DRI) in a 24x7 on-call rotation, monitoring service health and responding to incidents within SLA timelines.
- Automation & Deployment: Contribute to automation efforts and validate code functionality in non-production environments to ensure smooth deployments.
- Compliance & Security: Support compliance processes by verifying security, privacy, and accessibility standards during onboarding of new technologies.
- Continuous Learning: Stay current with industry trends and internal tools to improve reliability, performance, and observability at scale.
- Engineering Best Practices: Apply proven development and scaling practices to meet performance and customer requirements.
- Cross-Team Collaboration: Communicate effectively with engineering partners to align on goals and deliver user-centric solutions.
- Incident Response & Postmortems: Address complex live site issues, implement mitigations, and document learnings through postmortems.
Other:
- Embody our company's Culture and Values
Qualifications
- Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration
- OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration
- OR equivalent experience.
Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
- Candidates must have an active TS and be willing and eligible to upgrade to TS/SCI (with polygraph) or have an active TS/SCI and be willing and eligible to upgrade to TS/SCI (with polygraph). This role will require candidates to maintain the TS/SCI (with polygraph) clearance. Ability to meet Microsoft, customer and/or government security screening requirements are required pre-offer and post-hire for this role. Failure to maintain or obtain the appropriate clearance and/or customer screening requirements may result in employment action up to and including termination.
- Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
- Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter
- Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration
- OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration
- OR equivalent experience.
- 2+ years technical experience working with large-scale cloud or distributed systems.
- Demonstrated experience applying software engineering principles to production systems, including designing, building, or improving services and platforms.
- Proficiency in one or more programming languages such as C#, Go, Java, or Python, with the ability to develop and maintain production-quality code.
- Experience with automation that results in measurable improvements (e.g., reduced toil, fewer manual steps, improved system reliability).
- Experience with debugging and troubleshooting complex distributed systems in production environments.
- Ability to independently identify problems and implement solutions that improve system reliability and operational efficiency.
- Hands-on experience with CI/CD pipelines, testing, deployment, and reliability tooling.
Site Reliability Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Skills Required
- Master's in Computer Science/IT + 1+ year technical experience OR Bachelor's in Computer Science/IT + 2+ years technical experience OR equivalent experience
- Active Top Secret (TS) clearance and ability/willingness to obtain and maintain Top Secret/SCI (with polygraph)
- Pass Microsoft Cloud background check upon hire and every two years
- Serve as a Designated Responsible Individual (DRI) in a 24x7 on-call rotation and respond to incidents within SLA timelines
- Operate production services, perform incident response, implement mitigations, and conduct postmortems
- Experience working with large-scale cloud or distributed systems
- Apply software engineering principles to production systems and improve services/platforms
- Proficiency in one or more programming languages: C#, Go, Java, or Python
- Experience with automation that reduces toil and improves reliability
- Experience debugging and troubleshooting complex distributed systems in production
- Hands-on experience with CI/CD pipelines, testing, deployment, and reliability tooling
- Ability to independently identify problems and implement reliability and operational efficiency solutions
Microsoft Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.
-
Fair & Transparent Compensation — Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
-
Retirement Support — Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
-
Parental & Family Support — Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.
Microsoft Insights
What We Do
At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.







