Senior Software Engineer

Reposted 9 Days Ago
Be an Early Applicant
Redmond, WA, USA
In-Office
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Design, build, and operate globally distributed systems for Azure AI traffic shaping, routing, scheduling, capacity allocation, quota enforcement, and workload optimization. Lead architecture decisions, improve reliability and GPU utilization, develop observability and automation, resolve production issues, collaborate across Azure and OpenAI teams, mentor engineers, and influence platform strategy.
Summary Generated by Built In
Overview

At Microsoft, we’re a community of passionate innovators driven by curiosity and purpose. We collaborate to imagine what’s possible and accelerate our careers in a cloud-powered world where openness and innovation unlock limitless potential. 

Artificial Intelligence is central to Microsoft’s strategy—and the Azure AI Platform is leading the charge. As part of our team, you’ll contribute to cutting-edge projects that solve real-world challenges using transformative technologies. 

We are looking for a Senior Software Engineer to join our fast-paced, agile team at the core of Microsoft’s AI infrastructure. This team builds and operates the Next Generation Traffic Shaping, Scheduling, and Optimization Platform—a foundational control-plane service responsible for intelligent routing, capacity allocation, quota enforcement, and workload optimization for OpenAI models and other large-scale AI workloads across Azure. 

In this role, you will help define and build the systems that determine how AI traffic is routed and served globally. You will work on the Traffic Shaper, the critical platform that dynamically balances demand across thousands of AI endpoints, optimizes utilization of premium AI accelerators, and ensures reliable, low-latency experiences for Microsoft Copilot, Azure OpenAI Service, and strategic enterprise customers. 

You will design and develop large-scale distributed systems responsible for: 

  • Intelligent request routing and workload placement across global regions. 

  • Capacity management and allocation across models, tenants, and customer offerings. 

  • Real-time traffic shaping, prioritization, and load balancing. 

  • Quota management, throttling, and fairness enforcement. 

  • Fleet-wide optimization to maximize GPU utilization and service efficiency. 

  • High-fidelity telemetry, observability, and automated operational controls. 

  • Resiliency, failover, and incident mitigation mechanisms for mission-critical AI services. 

As a senior engineer, you will provide technical leadership across projects, drive architecture decisions, influence platform strategy, and partner closely with teams across Azure AI, OpenAI, Infrastructure, Networking, and Core AI Services. You will own services operating at global scale and help shape the future of Microsoft's AI infrastructure. 

astructure.

 

Impact 

This team sits on the critical path of Microsoft's AI strategy. The systems you build will directly influence how capacity is allocated, how traffic is routed, and how efficiently some of the world's largest AI workloads are served. Your work will help power Azure OpenAI Service, Microsoft Copilot experiences, and the next generation of AI products used by millions of customers worldwide. 


Responsibilities
  • Design and implement highly scalable, reliable, and performant distributed systems that power Azure AI traffic management and scheduling. 

  • Drive architecture and technical direction for routing, capacity allocation, and workload optimization services. 

  • Develop algorithms and control-plane systems that intelligently distribute AI workloads across global infrastructure while meeting latency, availability, and cost objectives. 

  • Collaborate with partner teams to onboard new models, regions, and customer offerings while maintaining operational excellence. 

  • Build observability and operational tooling to monitor service health, capacity utilization, traffic patterns, and customer impact in near real-time. 

  • Lead investigations into complex production issues and drive systematic improvements to reliability, resiliency, and platform automation. 

  • Identify and execute opportunities to improve fleet efficiency, reduce operational overhead, and optimize utilization of premium AI infrastructure. 

  • Mentor engineers, review designs and code, and raise engineering standards across the organization. 

  • Drive engineering excellence through testing, automation, documentation, and operational best practices. 

  • Influence long-term platform strategy and contribute to roadmap planning for next-generation AI infrastructure. 


Qualifications

Required: 

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Preferred: 

  • Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • 3+ years of experience in software engineering, preferably in distributed systems or cloud infrastructure. 

  • Strong coding skills in C#, Python, or Go. 

  • Experience with telemetry, metrics pipelines, or resource scheduling systems. 

  • Familiarity with cloud platforms (Azure, AWS, GCP) and container orchestration (Kubernetes, Service Fabric). 

  • Experience operating large-scale services with stringent reliability and availability requirements. 

  • Experience with distributed storage systems, telemetry platforms, and real-time data processing. 

  • Familiarity with AI/ML infrastructure, GPU fleets, inference serving systems, or large-scale model deployment platforms. 

  • Experience leading cross-team initiatives and driving projects from design through production deployment and operation. 


#AIPLATFORM, #AzureOpenAI, #llm, #distributedsystems, #MicrosoftAI, #AICoreServices  


Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.


Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor's degree in Computer Science or a related technical field, plus 4+ years of technical engineering experience with coding, or equivalent experience
  • Master's degree in Computer Science or a related technical field plus 6+ years of technical engineering experience, or bachelor's degree plus 8+ years, or equivalent experience
  • 3+ years of software engineering experience, preferably in distributed systems or cloud infrastructure
  • Strong coding skills in C#, Python, or Go
  • Experience with telemetry, metrics pipelines, or resource scheduling systems
  • Familiarity with cloud platforms such as Azure, AWS, or GCP
  • Familiarity with container orchestration technologies such as Kubernetes or Service Fabric
  • Experience operating large-scale services with stringent reliability and availability requirements
  • Experience with distributed storage systems, telemetry platforms, and real-time data processing
  • Familiarity with AI/ML infrastructure, GPU fleets, inference serving systems, or large-scale model deployment platforms
  • Experience leading cross-team initiatives and driving projects from design through production deployment and operation
  • Ability to pass the Microsoft Cloud Background Check and applicable customer or government security screenings

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
131K-260K Annually

inKind Logo inKind

Senior Software Engineer

eCommerce • Fintech • Food • Mobile • Social Impact
Remote or Hybrid
USA
170 Employees
160K-185K Annually

Airwallex Logo Airwallex

Senior Software Engineer

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office or Remote
Seattle, WA, USA
2300 Employees
180K-240K Annually

NinjaOne Logo NinjaOne

Senior Software Engineer

Information Technology • Productivity • Software • Infrastructure as a Service (IaaS)
Remote or Hybrid
18 Locations
2000 Employees
140K-200K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account