Principal AI Architect - M365 IC3 Team (Intelligent Conversation and Communications Cloud)

Reposted Yesterday
Be an Early Applicant
Redmond, WA, USA
In-Office
143K-304K Annually
Expert/Leader
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Architects and scales AI, LLM, RAG, and agent evaluation systems across the product lifecycle. Defines evaluation frameworks, quality metrics, dashboards, release gates, and monitoring workflows for nondeterministic agentic systems. Identifies model, retrieval, orchestration, UX, and product-quality failures; integrates evaluation into engineering pipelines; drives cross-functional architecture and technical strategy; and mentors engineers and applied scientists on reliable, responsible AI infrastructure.
Summary Generated by Built In
Overview

We are looking for a Principal AI Architect to help define, build, and scale the evaluation systems that shape the future of AI products. This role sits at the intersection of engineering, applied science, product architecture, and AI evaluation. The ideal candidate has deep technical judgment, strong systems thinking, hands-on experience evaluating LLMs, and the ability to translate emerging AI capabilities into reliable product experiences. 

This role is especially important for agentic AI systems, where product behavior is often nondeterministic, context-dependent, and difficult to evaluate with traditional testing alone. The person in this role will help teams build a deep understanding of how agents behave, where they succeed, where they fail, which gaps matter most, and where those gaps should be addressed: in prompts, tools, orchestration, retrieval, ranking, product UX, safety systems, or core code. 

The Principal AI Architect will make AI product quality measurable, actionable, and deeply integrated into how teams build and ship. They will help teams evaluate product direction before code is complete, validate quality before launch, and continuously measure performance after release. 

They will give teams confidence in how agentic systems behave, where nondeterminism creates risk, which gaps matter, and where to address them. Their work will ensure that evaluation is not an afterthought, but a core part of the product lifecycle, engineering system, and release decision process. 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Define the technical vision and architecture for AI, LLM, RAG, and agent evaluation systems across the product lifecycle.
  • Design evaluation frameworks that assess product quality before implementation is complete, during development, at launch, and post-ship.
  • Build an understanding of how nondeterministic agentic systems behave across tasks, users, contexts, tools, and product surfaces.
  • Identify behavioral gaps, failure modes, model limitations, retrieval issues, orchestration defects, and product-quality risks before they become customer-impacting issues.
  • Determine where issues should be addressed across the system: model behavior, prompts, tool use, search and retrieval, ranking, grounding, orchestration, UX, policy, telemetry, or product code.
  • Partner with engineering, applied science, and data science teams to bring ML, DS, LLM, RAG, and agent evaluation methods directly into product codebases and development workflows.
  • Integrate evals into build pipelines, release gates, experimentation systems, and engineering workflows so evaluation becomes a standard part of how products are built and shipped.
  • Make evaluation results easy to access, interpret, and act on through dashboards, scorecards, quality reports, and product-health views.
  • Build systems that connect product telemetry, offline evaluation, human judgment, automated evals, experimentation, RAG quality, agent behavior, and customer-quality signals.
  • Understand and evaluate enterprise search, RAG, grounding, indexing, ranking, permissions, freshness, and relevance systems for products such as Copilot.
  • Translate ambiguous product goals into measurable evaluation strategies, success criteria, timelines, and technical plans.
  • Drive architecture decisions across components, services, data pipelines, model interfaces, search systems, retrieval layers, evaluation harnesses, dashboards, and reporting systems.
  • Work with product leaders to prioritize evaluation investments and align them with product milestones and release decisions.
  • Mentor senior engineers and applied scientists on building reliable, scalable, and reusable evaluation infrastructure.
  • Stay current with LLM evaluation methods, agentic systems, RAG evaluation, benchmark design, prompt/model behavior, experimentation, and responsible AI practices.
  • In this role, you will help evaluate products before the code is fully ready, before launch, and after they ship.
  • You will work across product, engineering, applied science, and data science teams to bring rigorous LLM, RAG, agent, and AI evaluation practices into the product lifecycle.
  • You will help IC3 and partner teams understand whether AI systems are working as intended, where they fail, how they improve, and what it takes to ship them responsibly at scale. 

Qualifications
Required Qualifications:
  • Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research)
    • OR Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 4+ years related experience (e.g., statistics, predictive analytics, research)
    • OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 3+ years related experience (e.g., statistics, predictive analytics, research)
    • OR equivalent experience.
Preferred Qualifications:
  • Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 9+ years related experience (e.g., statistics, predictive analytics, research)
    • OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research) 
    • OR equivalent experience.
  • 5+ years experience creating publications (e.g., patents, libraries, peer-reviewed academic papers).
  • 2+ years experience presenting at conferences or other events in the outside research/industry community as an invited speaker.
  • 5+ years experience conducting research as part of a research program (in academic or industry settings).
  • 3+ years experience developing and deploying live production systems, as part of a product team.
  • 3+ years experience developing and deploying products or systems at multiple points in the product cycle from ideation to shipping.
  • Demonstrated experience evaluating LLMs, including designing eval datasets, defining quality metrics, analyzing model behavior, identifying failure modes, and using results to guide product or system improvements.
  • Understanding of modern AI systems, including LLMs, agentic systems, RAG, ML pipelines, evaluation methodology, experimentation, and product telemetry.
  • Ability to architect complex systems where multiple components, services, models, data flows, tools, retrieval systems, and product surfaces interact.
  • Experience evaluating or building nondeterministic AI systems where quality must be understood statistically, behaviorally, and through product impact.
  • Experience bringing ML, DS, LLM, or applied science concepts into production systems and product codebases.
  • Ability to integrate evaluation into engineering systems such as CI/CD, build pipelines, release gates, dashboards, and monitoring workflows.
  • Understanding of search, retrieval, grounding, relevance, ranking, and enterprise RAG concepts.
  • Ability to define technical strategy, product-quality metrics, milestones, and execution plans across teams.
  • Coding and technical design skills, with the ability to work directly in product codebases when needed.
  • Effective communication skills with the ability to influence engineers, scientists, product managers, and executives.
  • Track record of leading ambiguous, cross-functional technical initiatives from concept through delivery.
  • Experience evaluating LLM-powered products, agents, enterprise search, recommendation systems, or generative AI applications.
  • Experience with offline evals, online experimentation, human evaluation, red teaming, synthetic data, model monitoring, RAG evaluation, and agent behavior analysis.
  • Experience with evaluation dashboards, scorecards, quality reporting, product-health monitoring, or data visualization systems.
  • Familiarity with responsible AI, safety, reliability, privacy, security, permissions, compliance, and enterprise-readiness considerations for AI systems.
  • Experience building evaluation platforms, experimentation systems, model observability, agent evaluation infrastructure, or product-quality infrastructure.
  • Experience operating at principal, architect, or senior technical leadership level.
#LLM #Architect #EngineerScientist

Applied Sciences IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Bachelor's degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or a related field, plus 6 or more years of related experience; or equivalent experience.
  • Master's degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or a related field, plus 4 or more years of related experience; or equivalent experience.
  • Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or a related field, plus 3 or more years of related experience; or equivalent experience.
  • Master's degree in a related field plus 9 or more years of related experience, or doctorate plus 6 or more years of related experience.
  • Five or more years creating publications, such as patents, libraries, or peer-reviewed academic papers.
  • Two or more years presenting at conferences or industry events as an invited speaker.
  • Five or more years conducting research as part of an academic or industry research program.
  • Three or more years developing and deploying live production systems as part of a product team.
  • Three or more years developing and deploying products or systems across multiple product-cycle stages, from ideation through shipping.
  • Experience evaluating LLMs, including designing evaluation datasets, defining quality metrics, analyzing model behavior, identifying failure modes, and guiding product improvements.
  • Understanding of modern AI systems, including LLMs, agentic systems, RAG, machine learning pipelines, evaluation methodology, experimentation, and product telemetry.
  • Experience architecting complex systems involving multiple components, services, models, data flows, tools, retrieval systems, and product surfaces.
  • Experience evaluating or building nondeterministic AI systems using statistical, behavioral, and product-impact analysis.
  • Experience bringing machine learning, data science, LLM, or applied science concepts into production systems and product codebases.
  • Ability to integrate evaluation into CI/CD, build pipelines, release gates, dashboards, and monitoring workflows.
  • Understanding of search, retrieval, grounding, relevance, ranking, and enterprise RAG concepts.
  • Ability to define technical strategy, product-quality metrics, milestones, and execution plans across teams.
  • Coding and technical design skills, including the ability to work directly in product codebases.
  • Effective communication skills and ability to influence engineers, scientists, product managers, and executives.
  • Track record leading ambiguous, cross-functional technical initiatives from concept through delivery.
  • Experience evaluating LLM-powered products, agents, enterprise search, recommendation systems, or generative AI applications.
  • Experience with offline evaluations, online experimentation, human evaluation, red teaming, synthetic data, model monitoring, RAG evaluation, and agent behavior analysis.
  • Experience with evaluation dashboards, scorecards, quality reporting, product-health monitoring, or data visualization systems.
  • Familiarity with responsible AI, safety, reliability, privacy, security, permissions, compliance, and enterprise-readiness considerations.
  • Experience building evaluation platforms, experimentation systems, model observability, agent evaluation infrastructure, or product-quality infrastructure.
  • Experience operating at a principal, architect, or senior technical leadership level.

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Citizens Logo Citizens

Wealth Advisor - NW Pittsburgh, PA

Digital Media • Fintech • Information Technology • Machine Learning • Financial Services • Cybersecurity • Automation
In-Office or Remote
2 Locations
17000 Employees
105K-250K Annually

PwC Logo PwC

IT Audit/SOX - Senior Associate

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
Seattle, WA, USA
370000 Employees
77K-202K Annually

PwC Logo PwC

Front Office Strategy Consulting - Pharma Life Sciences Customer Analytics - Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
9 Locations
370000 Employees
99K-232K Annually

PwC Logo PwC

Oracle PMO - Senior Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
68 Locations
370000 Employees
124K-280K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account