Staff Machine Learning Engineer, Agent Eval Platform

Posted 5 Hours Ago
Be an Early Applicant
Mountain View, CA, USA
Hybrid
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
We're putting AI to work for people.
The Role
Build Moveworks’ agent evaluation and reward-modeling platform. Design configurable rubrics, deterministic validators, LLM judges, confidence scoring, and calibration against human-labeled trajectories. Analyze simulator-to-production divergence, maintain evaluation datasets, fine-tune smaller judge models, and develop process-level reward signals for optimizing prompts, tool selection, planning, retrieval, and routing. Establish versioned scenarios and reliable measurement standards for agent performance in stateful enterprise environments.
Summary Generated by Built In
Company Description

Who we are

Moveworks: the Agentic AI Assistant platform that empowers the entire workforce. 

Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt.

Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. 

In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work.

By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. Come join us!

ServiceNow: it all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.

Job Description

The Role

Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better?

That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.

This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

 

What you get to do in this role:

Judge design and calibration

  • A shared base judge with per-item rubrics expressed as configuration next to the dataset — so eval authors express intent, rather than forking a prompt per eval
  • Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy — was the clarifying question appropriate, was policy followed, was the path efficient
  • Scoring that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data
  • A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team — they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job
  • Fine-tuning a small judge model where an off-the-shelf one isn't good enough
  • Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction
  • Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit

Self-learning for the agent harness

This is where the pillar is headed, and a large part of why the seat exists.

  • A calibrated trajectory judge is, functionally, a reward model. Turning ours into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock
  • Using that signal to optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic
  • Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement
  • Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success

 

 

Qualifications

To be successful in this role you have:

  • 8+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used
  • Experience turning subjective human judgement into a measurement that holds up — one that other people, and ideally other models, can act on. This is the core of the job
  • Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain
  • Strong Python, and the discipline to ship production-grade code rather than notebooks
  • Ability to think and communicate clearly about complex problems — a large part of this job is convincing engineers that a number means what you say it means, and being right
  • A high degree of ownership and a bias toward shipping at startup pace
  • Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on

Experience in at least 3 of these:

  • LLM-as-judge or automated evaluation design, and calibrating it against human judgement
  • Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
  • Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly
  • Fine-tuning and evaluating small models: SFT, preference tuning, distillation
  • Reward modeling, RLHF/RLAIF, or process reward models
  • Agent trajectory analysis and step-level fault attribution
  • Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes

 

  

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. 

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license. 

Skills Required

  • 8+ years of experience in applied machine learning, data science, or ML-adjacent engineering
  • Track record of shipping and deploying work that was used in production
  • Experience converting subjective human judgment into reliable measurement systems
  • Strong applied machine learning fundamentals
  • Experience evaluating, prompting, and fine-tuning large language models
  • Strong Python programming skills and production-grade software engineering discipline
  • Ability to communicate clearly about complex technical problems
  • High degree of ownership and bias toward shipping quickly
  • Comfort working with ambiguity and exercising measurement judgment
  • Experience in at least three of: LLM-as-judge evaluation, human annotation programs, search or recommendation evaluation, small-model fine-tuning, reward modeling or RLHF/RLAIF, agent trajectory analysis, or prompt engineering

What the Team is Saying

Shanequa
Katya
Suzanne
Alexander
Jaime
Pat
Brady
Hasan
Jamil
Viviana
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
29,000 Employees
Year Founded: 2004

What We Do

As the AI platform for business transformation, we're putting AI to work across organizations — freeing people for work that matters. Making old tech work with new tech. Reaching across departments, from the front office to the back office and every office in between. Our ambition? To become the AI defining enterprise software company of the 21st century (or "AI DESCO21C," as we like to call it). With more than 8,400+ customers, we serve approximately 90% of the Fortune 500®, and we're proud to be a Fortune 100 Best Companies to Work For® and World's Most Admired Companies™. Explore your future career with us, visit www.careers.servicenow.com From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Why Work With Us

By joining ServiceNow, you are part of an ambitious team of change-makers who have a restless curiosity and a drive for ingenuity. We're committed to helping our people do their best work and live their best lives so we can fulfill our purpose together. At the fastest-growing enterprise software company, you can grow your career faster.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

ServiceNow Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

At ServiceNow, we lead with flexibility and trust. For some, home is the primary workplace. For those who come into a ServiceNow workplace, you are empowered to make team-guided and individual-led decisions on how and when you use the workplace.

Typical time on-site: Flexible
Company Office Image
HQSanta Clara, CA
Heredia
Ciudad de México
District of Columbia
Osaka
Aarhus, DK
Aarhus, DK
Company Office Image
Addison, TX
Amsterdam, North Holland
Atlanta, Georgia
Auckland, Auckland
Bangkok, Bangkok
Bengaluru, Karnataka
Bengaluru, Karnataka
Berlin, Berlin
Brasília, Federal District
Brisbane, Queensland
Brussels, BE
Cairo, Cairo Governorate
Canberra, Australian Capital Territory
Charlottesville, Virginia
Company Office Image
Chicago, IL
Deerfield, Illinois
Company Office Image
Denver, CO
Dubai, Dubai
Dublin, Leinster
Düsseldorf, Nordrhein-Westfalen
Frankfurt am Main, Hesse
Goteborg, Västra Götaland County
Gurugram, Haryana
Hamburg, Hamburg
Hanyang, Seoul
Helsinki, Uusimaa
Hong Kong, Hong Kong
Houston, TX
Hyderabad, Telangana
Issy-les-Moulineaux, Île-de-France
Johannesburg, Gauteng
Kirkland, WA
Lausanne, Vaud
Lille, Hauts de France
London, England
London, England
Madrid, Community of Madrid
Melbourne, Victoria
Milano, Lombardia
Milwaukee, WI
Company Office Image
Montréal, QC
Mumbai, Maharashtra
Munich, Bavaria
Company Office Image
New York, NY
Opfikon, Zürich
Orlando, FL
Oslo, Oslo
Perth, Western Australia
Petah Tikva, Central District
Company Office Image
Pleasanton, CA
Riyadh, Riyadh Province
Rome, Lazio
Company Office Image
San Diego, CA
San Francisco, Heredia
Company Office Image
San Francisco, CA
São Paulo, SP
Singapore, SG
Solna, Stockholm County
Sydney, New South Wales
Tokyo, Tokyo
Toronto, Ontario
Vancouver, British Columbia
Company Office Image
Vienna, VA
Vienna, AT
Company Office Image
Waltham, MA
Washington, DC
Wellington, Wellington
West Palm Beach, Florida
Learn more

Similar Jobs

ServiceNow Logo ServiceNow

Platform Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
San Diego, CA, USA
29000 Employees
181K-317K Annually

ServiceNow Logo ServiceNow

Staff Software Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Mountain View, CA, USA
29000 Employees

ServiceNow Logo ServiceNow

Tech Lead, Agent Eval Platform

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Mountain View, CA, USA
29000 Employees

ServiceNow Logo ServiceNow

Machine Learning Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Mountain View, CA, USA
29000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account