Principal Security Research Manager

Reposted 13 Days Ago
Be an Early Applicant
Redmond, WA, USA
In-Office
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
The Role
Lead design and implementation of end-to-end evaluation platforms for agentic security workflows: build benchmarks, deterministic validators, graders, telemetry, CI/CD integration, and production feedback loops to measure safety, correctness, cost, latency, and human-review effort. Partner with engineering and product teams to detect regressions, compare configurations, and provide release-readiness evidence and reusable playbooks.
Summary Generated by Built In
Overview

We are building the evaluation backbone for safe, reliable, and efficient agentic engineering in Microsoft Security. This team will create systems that determine when an AI agent, model, prompt, tool, memory strategy, or orchestration pattern is ready to be used in production security and engineering workflows. The role is ideal for engineers who can operate end to end: understand the workflow, design the benchmark, build the harness, implement validators and graders, run experiments, analyze quality and cost tradeoffs, connect results to production feedback, and help teams make evidence-based release decisions. 

Why this role matters: Microsoft Security is moving toward agentic engineering systems for security triage, remediation, repo readiness, and scan-to-verified-closure workflows. Evals are the trust system for that shift. They help decide whether autonomy can safely expand, whether a release should stop, and which configuration achieves the required quality, safety, reliability, latency, and cost bar with the lowest practical human-review burden. 

Role mission 

As a Principal Security Research Manager on the AI Evaluation Systems team, you will build the common evaluation platform and methodology used by MSec agent programs. You will work across evaluation design, platform implementation, test infrastructure, telemetry, measurement, security workflow understanding, and production learning. Your work will make agentic systems measurable, reproducible, governable, and continuously improving. 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
 
  • Design and build end-to-end evaluation harnesses for agentic security and engineering workflows, including triage, remediation, repo readiness, escalation, tool use, and scan-to-verified-closure paths. 

  • Create representative benchmark suites and golden datasets that include normal, edge, adversarial, failure-recovery, regression, and production-derived cases. 

  • Implement deterministic validators, automated graders, trace analyzers, result stores, comparison views, and workflow adapters that make evaluations repeatable and actionable. 

  • Measure task success, correctness, safety and policy compliance, failure recovery, latency, tool-call behavior, token usage, total cost per successful outcome, and human-review effort. 

  • Compare models, prompts, tools, memory strategies, policies, and orchestration patterns under consistent conditions and help teams understand quality-versus-efficiency tradeoffs. 

  • Integrate evaluations into engineering workflows, CI/CD, release gates, and decision processes so material agent changes are supported by reproducible evidence before production rollout. 

  • Connect offline evaluation results with production feedback, including accepted and rejected outputs, human overrides, incidents, rollbacks, escaped defects, and customer or service-health signals. 

  • Detect regressions, benchmark drift, evaluator miscalibration, and cases where eval scores improve while real-world outcomes do not. 

  • Partner with agent builders, product teams, security engineers, data scientists, program managers, and leadership to turn evaluation results into release recommendations and autonomy-boundary decisions. 

  • Convert learnings into reusable paved paths: playbooks, templates, onboarding guides, dashboards, scorecards, and reference implementations that can scale across MSec. 

What you will build 

  • Versioned benchmark suites for security triage, remediation, repo readiness, tool use, escalation, and end-to-end agentic workflows. 

  • Evaluation runners and harnesses that can replay tasks, capture traces, evaluate outputs, and compare multiple agent configurations. 

  • Deterministic validation checks for code, policy, security, provenance, ownership, deployment constraints, and workflow-specific correctness. 

  • Automated and human-in-the-loop grading pipelines with calibration, sampling, and confidence thresholds. 

  • Pareto-style scorecards that show tradeoffs across quality, risk, latency, tokens, cost, and human-review burden. 

  • Telemetry and production feedback loops that continuously expand benchmark coverage and keep offline evaluation anchored to real-world outcomes. 

  • Release-readiness gates and evidence packages that help leaders and product teams decide whether to scale, stop, or redesign an agent pattern. 


Qualifications

Required Qualifications:

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 3+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
    • OR Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 4+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
    • OR Bachelor's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 6+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
    • OR equivalent experience.
  • 1+ year(s) people management experience.

Preferred Qualifications: 

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 5+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
    • OR Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 8+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
    • OR Bachelor's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 12+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
    • OR equivalent experience.
  • Proven software engineering experience building production systems, developer platforms, test infrastructure, automation frameworks, data pipelines, quality systems, or reliability tooling. 
  • Ability to design and implement evaluation systems end to end, including task definition, dataset creation, harness implementation, scoring, analysis, and operational integration. 
  • Demonstrated coding, debugging, system design, and operational excellence skills. 
  • Experience working with structured data, logs, traces, metrics, APIs, automation workflows, and engineering telemetry. 
  • Experience with LLMs, AI agents, model evaluation, prompt/tool orchestration, automated grading, evaluation harnesses, or benchmark design. 
  • Experience with experimentation, statistical confidence, evaluator calibration, regression analysis, human-review protocols, or quality measurement systems. 
  • Experience with security engineering, vulnerability management, SAST/SCA, SARIF, remediation workflows, secure development lifecycle, or compliance-sensitive systems. 

Security Research M5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Skills Required

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 3+ years relevant experience; OR Master's AND 4+ years; OR Bachelor's AND 6+ years; OR equivalent experience.
  • 1+ year people management experience.
  • Proven software engineering experience building production systems, developer platforms, test infrastructure, automation frameworks, data pipelines, or reliability tooling.
  • Ability to design and implement evaluation systems end to end: task definition, dataset creation, harness implementation, scoring, analysis, and operational integration.
  • Demonstrated coding, debugging, system design, and operational excellence skills.
  • Experience with structured data, logs, traces, metrics, APIs, automation workflows, and engineering telemetry.
  • Experience with LLMs, AI agents, model evaluation, prompt/tool orchestration, automated grading, evaluation harnesses, or benchmark design.
  • Experience with experimentation, statistical confidence, evaluator calibration, regression analysis, and human-review protocols.
  • Experience with security engineering, vulnerability management, SAST/SCA, SARIF, remediation workflows, secure development lifecycle, or compliance-sensitive systems.

Microsoft Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.

  • Fair & Transparent Compensation Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
  • Retirement Support Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
  • Parental & Family Support Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.

Microsoft Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redmond, WA
206,870 Employees
Year Founded: 1975

What We Do

At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.

Similar Jobs

Coursera + Udemy  Logo Coursera + Udemy

Fp&a Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
United States
1500 Employees
111K-162K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Lead Supervisor I

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
North Bend, WA, USA
16000 Employees
17-28 Hourly

Citizens Logo Citizens

Wealth Advisor - Lansdale, PA

Digital Media • Fintech • Information Technology • Machine Learning • Financial Services • Cybersecurity • Automation
In-Office or Remote
2 Locations
17000 Employees
105K-250K Annually

Pfizer Logo Pfizer

Research & Development Rotational Program Associate

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
5 Locations
121990 Employees
60K-100K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account