Responsibilities
The team is looking for an experienced engineer to:
Lead the design and development of large-scale distributed systems that manage resource allocation, workload execution, and service protection.
Define and drive technical strategy for platform capabilities that improve reliability, efficiency, scalability, and operational excellence.
Build intelligent, signal-driven automation using telemetry, health indicators, and real-time platform insights.
Develop solutions that balance customer experience, infrastructure utilization, operational cost, and service performance.
Drive innovations that proactively identify, mitigate, and prevent service disruptions.
Partner with engineering teams across Microsoft 365, Copilot, Azure, and infrastructure organizations to deliver end-to-end platform solutions.
Influence architecture, design, and engineering best practices across multiple teams.
Mentor engineers and contribute to a culture of technical excellence and continuous improvement.
Help shape how future AI-powered services are scaled, managed, protected, and optimized across Microsoft 365.
Qualifications
Required Qualifications:
- Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR equivalent experience.
- Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR equivalent experience.
- Deep understanding of distributed systems, concurrency, reliability engineering, scalability, and platform architecture.
- Experience developing and operating highly available backend services at scale.
- Demonstrated ability to lead technically complex initiatives across multiple teams and organizations.
- Solid problem-solving skills involving system performance, resiliency, resource management, and operational excellence.
- Demonstrated ability to effectively leverage AI-assisted engineering tools and autonomous coding agents to improve software development productivity, quality, and operational effectiveness.
- Ability to critically evaluate, validate, and refine AI-generated code, designs, tests, diagnostics, and recommendations while maintaining full engineering ownership and accountability.
- Experience building infrastructure platforms, resource management systems, scheduling systems, load balancing systems, storage platforms, or large-scale backend services.
- Solid understanding of workload management, traffic engineering, fault tolerance, admission control, and capacity planning.
- Experience using telemetry, monitoring, and service health signals to drive automated operational decisions.
- Experience supporting services with rapidly changing demand patterns and large-scale customer workloads.
- Hands-on experience with cloud-native architectures and distributed platforms such as Azure or similar cloud environments.
Experience with:
- Microservices and service-oriented architectures
- Event-driven systems
- Large-scale telemetry systems
- Containerized and cloud-native environments
- Experience building or supporting AI-powered services and high-throughput systems.
- Familiarity with AI workload characteristics, including bursty traffic, latency sensitivity, resource contention, and dynamic scaling requirements.
- Experience integrating AI-assisted engineering workflows into software development, testing, debugging, code review, and operational processes.
- Proven ability to drive technical alignment and influence engineering decisions across organizational boundaries.
- Solid communication skills with the ability to explain complex technical concepts to diverse audiences.
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Skills Required
- Bachelor's degree in Computer Science or a related technical field, or equivalent experience
- 6+ years of technical engineering experience with coding in C, C++, C#, Java, JavaScript, Python, or similar languages
- Master's degree in Computer Science or a related technical field and 8+ years of technical engineering experience, or bachelor's degree and 12+ years, or equivalent experience
- Deep understanding of distributed systems, concurrency, reliability engineering, scalability, and platform architecture
- Experience developing and operating highly available backend services at scale
- Ability to lead technically complex initiatives across multiple teams and organizations
- Experience with infrastructure platforms, resource management systems, scheduling systems, load balancing systems, storage platforms, or large-scale backend services
- Understanding of workload management, traffic engineering, fault tolerance, admission control, and capacity planning
- Experience using telemetry, monitoring, and service health signals to drive automated operational decisions
- Experience with Azure or similar cloud-native distributed platforms
- Experience with microservices, service-oriented architectures, event-driven systems, telemetry systems, and containerized environments
- Experience building or supporting AI-powered services and high-throughput systems
- Ability to use, evaluate, validate, and refine AI-assisted engineering tools and autonomous coding agents
- Ability to drive technical alignment and influence engineering decisions across organizational boundaries
- Strong communication skills for explaining complex technical concepts to diverse audiences
Microsoft Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.
-
Fair & Transparent Compensation — Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
-
Retirement Support — Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
-
Parental & Family Support — Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.
Microsoft Insights
What We Do
At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.








