Principal Software Engineer - AKS Flex: High Scale

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Lead architecture and development for AKS Flex, scaling etcd and Kubernetes control planes for distributed node fleets. Design remote-node provisioning, lifecycle automation, Kubernetes-native APIs, secure bootstrap, upgrades, and reliability systems across Azure, on-premises, edge, and other clouds. Drive performance testing, production operations, open-source collaboration, technical strategy, cross-team alignment, and mentorship while solving distributed-systems, storage, security, and scalability challenges.
Summary Generated by Built In
Overview

Imagine working at the forefront of the cloud industry, redefining where Kubernetes can run. The AKS Flex team is building the next generation of Azure Kubernetes Service: a single AKS control plane that can manage worker nodes running anywhere, whether in other Azure regions, on-premises data centers, bare-metal fleets, edge sites, or other clouds.

We are hiring a Principal Software Engineer to scale this vision. You will lead the scaling of etcd and the Kubernetes control plane so a single cluster can reliably manage far larger, more geographically distributed, and more heterogeneous node fleets than ever before. You will also build the platform behind flex nodes for AKS and Unbounded, our open-source project for joining any Linux machine to a Kubernetes cluster, including provisioning, lifecycle, and the Kubernetes-native APIs that make remote machines manageable at scale.

You will solve deeply technical distributed-systems problems in control plane performance, consistency and durability, node provisioning and lifecycle, security, and reliable production operations. Your work will directly unlock capacity for GPU-hungry AI workloads, data-residency-constrained customers, and organizations that want one operating model for all of their compute.

This is a unique opportunity to define the architecture of a new product area from its early days, represent Microsoft in the open-source community, and shape what hyperscale managed Kubernetes looks like without boundaries.

Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities

Technical leadership

  • Own the architecture for control plane scalability and the flex nodes and Unbounded platform, and help define the multi-year technical roadmap for AKS Flex in partnership with product and engineering leadership.
  • Lead cross-organization design efforts spanning AKS, Azure compute, storage, networking, and identity, driving alignment on interfaces, tradeoffs, and sequencing.
  • Make and defend key architectural decisions around correctness, durability, security, scale limits, and failure modes, and establish the design review and quality bar for the area.
  • Represent Microsoft in upstream communities (such as Kubernetes SIG Scalability, SIG etcd, SIG API Machinery, and the Unbounded project), driving proposals and building consensus with external maintainers.
  • Mentor and grow senior engineers, and raise the technical capability of the broader team.

Control plane and etcd scalability

  • Push the limits of cluster scale by improving the performance, reliability, and efficiency of etcd and the Kubernetes API server under large node counts, high object churn, and heavy watch traffic.
  • Analyze and optimize control plane bottlenecks such as write amplification, compaction and defragmentation, watch fan-out, list/relist storms, request prioritization, and storage and memory pressure.
  • Lead the design and validation of architectural changes to the control plane storage layer, including approaches to partitioning, offloading, and caching state, while preserving correctness and durability guarantees.
  • Build scale and performance testing infrastructure, benchmarks, and SLOs that let us measure control plane limits and prevent regressions before they reach customers.
  • Build automation for safe etcd operations at fleet scale: backup and restore, member replacement, quorum recovery, upgrades, and repair.

Flex nodes and Unbounded platform

  • Design and build the systems that let worker nodes in other Azure regions, on-premises data centers, edge locations, and other clouds securely join an AKS cluster and behave as standard node pools.
  • Build node provisioning and day-2 lifecycle automation across diverse substrates: SSH-based bootstrap, cloud API provisioning in response to unschedulable pods, and bare-metal PXE boot with BMC power management and TPM-based attestation.
  • Design Kubernetes-native APIs, controllers, and CRDs that make remote machines declaratively manageable with standard tooling, and that scale efficiently as fleets grow.
  • Solve the control plane challenges of remote nodes: high and variable latency, intermittent connectivity, node heartbeats and status churn, node identity and authentication, secure bootstrap, and safe upgrades across heterogeneous hardware including GPUs.
  • Contribute to and help lead the Unbounded open-source project, including design discussions, code reviews, and community engagement.

Engineering excellence

  • Participate in production operations and on-call rotations, turning operational learnings into durable engineering improvements.
  • Partner across AKS and Azure teams to deliver secure, compliant, and resilient infrastructure.
  • Drive engineering culture through technical design, code reviews, and setting standards for reliability and operational excellence across the area.


Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, Go, Rust, C++, C#, Java, or Python
    • OR equivalent experience

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications:

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, Go, Rust, C++, C#, Java, or Python
    • OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, Go, Rust, C++, C#, Java, or Python
    • OR equivalent experience.
  • Deep understanding of etcd internals, including Raft consensus, MVCC, the watch mechanism, compaction and defragmentation, and the bbolt storage backend, as well as experience operating etcd at scale.
  • Deep understanding of storage systems and databases, including storage engines (B-trees, LSM trees), write-ahead logging, replication, consistency models, transactions, and performance tuning.
  • Experience with Kubernetes control plane internals: API server, watch cache, controller patterns, API priority and fairness, and scalability limits.
  • Experience with performance engineering: profiling, load testing, capacity modeling, and diagnosing latency and throughput issues in production.
#azurecorejobs

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor's degree in Computer Science or a related technical field, or equivalent experience
  • 6+ years of technical engineering experience with coding in Go, Rust, C++, C#, Java, Python, or similar languages
  • Ability to pass the Microsoft Cloud Background Check and applicable customer or government security screening requirements
  • Master's degree in Computer Science or a related technical field and 8+ years of technical engineering experience, or bachelor's degree and 12+ years of experience, or equivalent experience
  • Deep understanding of etcd internals, including Raft consensus, MVCC, watches, compaction, defragmentation, bbolt, and operating etcd at scale
  • Deep understanding of storage systems and databases, including B-trees, LSM trees, write-ahead logging, replication, consistency models, transactions, and performance tuning
  • Experience with Kubernetes control plane internals, including the API server, watch cache, controller patterns, API priority and fairness, and scalability limits
  • Experience with performance engineering, profiling, load testing, capacity modeling, and diagnosing production latency and throughput issues

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation — Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support — Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support — Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Superhuman Logo Superhuman

Counsel

Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
Remote or Hybrid
United States
1500 Employees
230K-352K Annually

GitLab Logo GitLab

Senior Professional Services Engineer - PubSec - US Only

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
United States
2500 Employees
136K-230K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Systems Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
Remote
California, USA
85422 Employees
203K-420K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Enterprise Architect

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office or Remote
4 Locations
85422 Employees
167K-391K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account