Principal Software Engineer - AI Systems Observability

Posted Yesterday
Hiring Remotely in United States
Remote
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Lead the architecture and hands-on development of agentic observability capabilities for Kubernetes and Linux. Build production-quality software, interpret systems telemetry using eBPF, mentor engineers, coordinate cross-team delivery, support open-source initiatives, and participate in on-call operations. The role focuses on safe, observable, and evaluable AI-agent workflows, distributed systems reliability, and technical leadership across engineering teams.
Summary Generated by Built In
Overview

Help shape how intelligent agents understand and operate the systems that power the cloud. On the Azure Core Upstream Observability team, you will build and contribute to open-source technologies that make Kubernetes and Linux environments easier to understand, diagnose, and operate. You will work with engineers, product managers, partner teams, and upstream communities to advance reliable, secure, and efficient observability capabilities. You will help the team turn emerging agentic-system opportunities into practical capabilities grounded in rich systems telemetry.

As a Principal Software Engineer, you will set technical direction and lead the design and delivery of agentic observability capabilities across Kubernetes and Linux. You will turn evolving customer, partner, and operational needs into coherent architectures and production-quality solutions while remaining hands-on in implementation, review, and operational excellence. You will connect work across organizational boundaries and support engineers and teams in delivering dependable, maintainable capabilities with measurable outcomes.
This opportunity will allow you to deepen your proficiency in safe, observable, and evaluable agentic systems. Expand your technical leadership across Kubernetes, Linux, and cloud-scale distributed systems. Strengthen your ability to shape strategy and support engineering outcomes across teams and open-source communities.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Partner with appropriate stakeholders to determine user requirements for agentic observability scenarios across Kubernetes and Linux environments.
  • Lead identification of dependencies and development of design documents for products, applications, services, or platforms that support safe, observable, and evaluable agentic workflows.
  • Lead by example and mentor others to produce extensible and maintainable code used across products, including hands-on implementation, review, testing, debugging, and maintenance.
  • Use proficiency in cross-product features to work cooperatively with appropriate stakeholders, including product managers, and drive project plans, release plans, and work items across multiple groups.
  • Apply Kubernetes and Linux proficiency to collect, correlate, and interpret systems signals, using Linux kernel and extended Berkeley Packet Filter (eBPF) techniques when they are the appropriate engineering choice.
  • Be responsible as a Designated Responsible Individual (DRI), support engineers across products and solutions, and participate in on-call work to monitor systems, products, and services for degradation, downtime, or interruptions.
  • Proactively seek new knowledge and adapt to trends, technical solutions, and patterns that improve availability, reliability, efficiency, observability, and performance while supporting consistent monitoring and operations at scale and sharing knowledge with other engineers. 

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, Go, Python, or Rust
    • OR equivalent experience.

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications:

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, Go, Python, or Rust
    • OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, Go, Python, or Rust
    • OR equivalent experience.
  • Experience designing and operating at least one production Kubernetes/Linux observability or distributed-systems capability, including telemetry, incident response, and rollout or rollback.
  • Experience builing or operating at least one tool-using AI-agent system in production or a production-like environment, with a documented evaluation set and auditable approval or safety controls. 
  • Experience leading adoption of a shared technical architecture by at least two engineering teams, or authored a design/change accepted by an upstream open-source project.  

#azurecorejobs


Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor's degree in Computer Science or a related technical field and 6+ years of technical engineering experience with coding in C, C++, Go, Python, or Rust, or equivalent experience
  • Ability to pass the Microsoft Cloud Background Check upon hire or transfer and every two years thereafter
  • Master's degree in Computer Science or a related technical field and 8+ years of technical engineering experience, or bachelor's degree and 12+ years of experience, or equivalent experience
  • Experience designing and operating a production Kubernetes/Linux observability or distributed-systems capability, including telemetry, incident response, and rollout or rollback
  • Experience building or operating a production or production-like tool-using AI-agent system with a documented evaluation set and auditable approval or safety controls
  • Experience leading adoption of a shared technical architecture by at least two engineering teams or authoring a design/change accepted by an upstream open-source project

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation — Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support — Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support — Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

ServiceNow Logo ServiceNow

Director, Life Sciences GTM, Pharma

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Raleigh, NC, USA
29000 Employees

BECU Logo BECU

Manager Procurement Operations

Fintech • Financial Services
Remote or Hybrid
3 Locations
3000 Employees
102K-189K Annually

DigitalOcean Logo DigitalOcean

Product Manager

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
In-Office or Remote
San Francisco, CA, USA
1400 Employees
186K-280K Annually

DigitalOcean Logo DigitalOcean

Product Manager

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
In-Office or Remote
Austin, TX, USA
1400 Employees
186K-280K Annually

Similar Companies Hiring

Ford Energy Thumbnail
Automotive • Software • Energy • Utilities • Manufacturing • Renewable Energy
US
55 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account