Software Engineer II

Posted 12 Hours Ago
Be an Early Applicant
Redmond, WA, USA
In-Office
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Design, develop, test, and operate distributed AI infrastructure services supporting large-scale training and inference. Build C# control-plane services on Service Fabric or Kubernetes, improve reliability, efficiency, security, and latency, and provide on-call operational support. Collaborate with engineering, product, research, and data science teams while developing expertise in AI/ML concepts. Contribute technical leadership, rigorous testing, data-driven troubleshooting, and high-quality service design for the GPU scheduling subsystem.
Summary Generated by Built In
Overview
The AI Infrastructure team is responsible for building and operating the large-scale, reliable, and efficient GPU-based clustering infrastructure that powers Microsoft’s AI/ML ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Azure OpenAI Service, M365 Copilot & Copilot Tuning, GitHub Copilot, Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models; as well as the mature Azure ML Services, which provide data scientists and developers a rich experience for defining, training, fine-tuning, deploying, monitoring, and consuming machine learning models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads and business groups across the company. We engage directly with some of the major internal research and applied AI/ML groups using these services, including Microsoft Research, M365, Microsoft Security, and the Bing WebXT team.
 
The AI Infra team is looking for a talented Software Engineer II, with initial focus on the Scheduler subsystem. The scheduler is the “brains” of the AI Infra control plane. It governs access to the GPU, NPU and CPU capacity of the platform according to a complex system of workload preference rules, placement constraints, optimization objectives, and dynamically interacting policies aimed to maximize hardware utilization and fulfill greatly varying needs of users and the AI platform partner services in terms of workload types, prioritization, and capacity targeting flexibility. The scheduler’s set of capabilities is broad and ambitions. It manages quota, capacity reservations, SLA tiers, preemption, auto-scaling, and a wide range of configurable policies. It is both a workload-aware and topology-aware scheduler (down to the level of cluster racks and nodes). Global scheduling is a distinctive major feature that overcomes the regional segmentation of the Azure compute fleet by treating the GPU capacity as a single global virtual pool, which greatly increases capacity availability and utilization for major classes of AI/ML workload. We have achieved this capability by avoiding a significant global single point of failure, based on regional instances of the scheduler service interacting via peer-to-peer protocols for sharing capacity inventory and coordinating handoff of jobs for scheduling. Our system manages significant amount of GPU capacity even outside Azure datacenters, through a unified model and operational process and highly generalized, flexible workload scheduling capabilities.
 
To be able to manage the inherent complexity of the Scheduler subsystem and enable it to meet the stringent expectations of high service reliability, availability, and throughput, we emphasize rigorous engineering, utmost precision and quality, and strong ownership—from feature design to livesite. Quality mindset, attention to detail, development process rigor, and data-driven design and problem-solving skills are key for success in our mission-critical control plane space. We enjoy great creative freedom and thrive on the capacity management and workload scheduling challenges posed to us by the various partner teams.
 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Work on the design and development of the core AI Infrastructure distributed and in-cluster services that support large scale AI training and inferencing.
  • Develop, test, and maintain control plane services written in C#, hosted on Service Fabric or Kubernetes (AKS) clusters.
  • Enhance systems and applications to ensure high stability, efficiency and maintainability, low latency, tight cloud security.
  • Provide operational support and DRI (on-call) responsibilities for the service.
  • Develop and foster a deep understanding of the AI/ML concepts, use cases, and relevant services used by our customers. Be an AI-first developer, making productive use of the available tools and actively engaging in experimentation and learning.
  • Collaborate closely with service engineers, product managers, and internal applied research and data science teams within Microsoft to build better solutions together.
  • Provide vision, expertise, and technical leadership to other team members.
  • Embody our culture and values
 

Qualifications

Required Qualifications: 

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience. 

Other Requirements:

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:

Microsoft Cloud Background Check:

- This position will be required to pass the Microsoft background and Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred Qualifications: 

  • Hands-on (devops) experience with larger-scale, high-availability cloud services at the PaaS or IaaS level, based on microservices architecture, ideally related to AI infrastructure or workload hosting
  • Proficiency with use of complex data structures and algorithms, preferably in the setting of a resource allocator/scheduler, workflow/execution orchestration engine, database engine, or similar
  • Proficiency and thoroughness in unit testing and testability techniques
  • Agentic development skills
  • Experience with building and operating “stateful” and critical control plane services; handling challenges with data size and data partitioning; advanced use of a NoSQL cloud database
  • Service reliability and fundamentals engineering; instrumentation for KPIs or performance analysis; demonstrated service and code quality mindset
  • Applied knowledge of Kubernetes: service model, workload packaging and deployment, programmatic extensibility (CRDs, operators); or equivalent knowledge of Service Fabric; experience with any service mesh
  • Data-driven design and troubleshooting and data analytics skills, ideally with Kusto
#AIINFRA
 

Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor’s degree in Computer Science or a related technical field
  • 2+ years of technical engineering experience with coding in languages including C, C++, C#, Java, JavaScript, or Python
  • Equivalent practical experience may substitute for the bachelor’s degree and experience requirement
  • Ability to meet Microsoft, customer, and/or government security screening requirements
  • Pass the Microsoft background and Microsoft Cloud background checks upon hire or transfer and every two years thereafter
  • Hands-on DevOps experience with large-scale, highly available PaaS or IaaS cloud services
  • Experience with microservices architecture, ideally in AI infrastructure or workload hosting
  • Proficiency with complex data structures and algorithms, preferably in resource allocation, scheduling, orchestration, or database systems
  • Proficiency in unit testing and testability techniques
  • Agentic development skills
  • Experience building and operating stateful, critical control-plane services
  • Experience handling data size and data partitioning challenges
  • Advanced use of a NoSQL cloud database
  • Service reliability, instrumentation, KPI measurement, performance analysis, and code quality experience
  • Applied knowledge of Kubernetes, including service models, workload packaging, deployment, CRDs, and operators, or equivalent Service Fabric knowledge
  • Experience with a service mesh
  • Data-driven design, troubleshooting, and data analytics skills, ideally using Kusto

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Expedia Group Logo Expedia Group

Software Engineer

AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Hybrid
Seattle, WA, USA
16000 Employees
119K-191K Annually

Redfin Logo Redfin

Software Engineer

Fintech • Real Estate • PropTech
Remote or Hybrid
Seattle, WA, USA
5800 Employees
139K-170K Annually

CrowdStrike Logo CrowdStrike

Engineer II, Software Assurance, Product Security (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
100K-145K Annually

Bestow Logo Bestow

Software Engineer

Big Data • Fintech • Information Technology • Insurance • Software
Remote or Hybrid
US
160 Employees
125K-146K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account