AI Policy Generalist - Seattle Onsite

Posted 2 Days Ago
Be an Early Applicant
Seattle, WA, USA
In-Office
45-55 Hourly
Entry level
Edtech • Enterprise Web • HR Tech • Software
Handshake is the number one site for college students to find a job.
The Role
Evaluate AI model requests and responses against customer policies, taxonomies, and rubrics. Make nuanced classifications, write evidence-based rationales, identify policy gaps, participate in calibration discussions, incorporate feedback, and help improve evaluation frameworks. The role requires consistent reasoning across sensitive subject matter, including self-harm, violence, sexual content, abuse, and discrimination.
Summary Generated by Built In
About Handshake

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.

Why join Handshake now:

  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel

  • Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world’s top educational institutions

  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders

  • Build a massive, fast-growing business with billions in revenue

About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

As an AI Policy Generalist, you will turn complex customer policies into consistent, well-reasoned evaluations of AI model behavior.

You will read user requests, model responses, and relevant conversation history, then determine which policy category best applies. The most interesting cases will not have obvious answers. Two examples may look almost identical until a single word, contextual detail, or difference in intent changes the correct classification.

We are looking for people who enjoy splitting hairs in a healthy way. You form clear opinions, explain precisely why two cases should be treated differently, challenge interpretations respectfully, and change your mind when better evidence emerges. You understand that productive disagreement is not about winning an argument. It is how a team finds the most accurate and consistent interpretation.

This is not rote annotation. Policies cannot anticipate every possible edge case, and good evaluators do not apply them mechanically. You will balance the policy’s text and intent with customer expectations, conversation context, precedent, and team calibration.

The subject matter will vary. One project may involve distinguishing benign assistance from meaningful facilitation of harm. Another may require evaluating whether an interaction reflects ordinary emotional support or unhealthy reliance. A third may focus on nuanced boundaries within sexual-safety policy. Success requires learning each customer’s framework on its own terms rather than carrying assumptions from one domain into another.

What You Will Do
  • Learn new customer policies, definitions, taxonomies, and evaluation rubrics quickly

  • Evaluate user requests and AI model responses within the full relevant conversation context

  • Distinguish between closely related labels, severity levels, and policy boundaries

  • Select the most defensible classification when a case is genuinely ambiguous

  • Write concise, evidence-based rationales that cite relevant policy language and conversation details

  • Identify policy gaps, contradictions, unclear definitions, and emerging edge cases

  • Raise thoughtful questions when existing guidance does not resolve a case

  • Participate actively in calibration discussions with evaluators, project leads, policy teams, and researchers

  • Challenge interpretations respectfully and update your judgment when new guidance or stronger reasoning emerges

  • Apply customer policy consistently without substituting personal beliefs for the policy standard

  • Maintain accuracy and attention to detail across repeated evaluations

  • Incorporate feedback quickly and apply clarified guidance to future work

  • Help improve evaluation frameworks, examples, decision rules, and quality standards

  • Move effectively between projects covering different policy domains and customer needs

You May Be a Fit If
  • You enjoy making precise distinctions between cases that other people might consider equivalent

  • You notice when one word, contextual detail, or change in intent materially affects the answer

  • You can hold a strong opinion without becoming attached to being right

  • You explain judgment calls clearly enough that another person can audit your reasoning

  • You ask productive questions when a policy is ambiguous instead of guessing or forcing certainty

  • You can separate your personal views from the standard a customer has asked you to apply

  • You are comfortable discussing disagreement directly, respectfully, and without making it personal

  • You can follow the letter of a policy while also understanding its purpose and underlying logic

  • You remain careful and consistent during repetitive, feedback-heavy work

  • You learn unfamiliar subject matter quickly and know when additional context is needed

  • You are intellectually curious, self-directed, and comfortable working in a fast-changing environment

  • You communicate clearly and precisely in writing

  • You treat sensitive information and difficult subject matter with maturity and sound judgment

Strong candidates may come from quality assurance, research, editing, law, teaching, operations, trust and safety, content moderation, social science, policy, investigations, compliance, customer support, or other fields that require careful interpretation and defensible decision-making. We care more about how you reason than where you learned to reason.

Nice to Have
  • Experience evaluating or comparing outputs from ChatGPT, Claude, Gemini, or other language models in a professional capacity

  • Prior work in AI evaluation, data annotation, RLHF, model quality, trust and safety, policy operations, or content moderation

  • Experience applying detailed rubrics, taxonomies, regulatory language, editorial standards, or quality frameworks

  • Familiarity with calibration sessions, inter-rater agreement, quality audits, or adjudication workflows

  • Experience writing policy guidance, decision trees, evaluation examples, or structured rationales

  • Comfort working with long conversations, incomplete context, and conflicting evidence

  • Familiarity with AI safety, responsible AI, or the ways language models can assist, mislead, or cause harm

Prior AI evaluation experience is helpful, but it is not required.

Sensitive-Content Notice

This role involves regular and deliberate engagement with sensitive material. Depending on the project, evaluations may include sexual content, emotional distress, self-harm, suicide, violence, weapons, abuse, exploitation, discrimination, and other potentially disturbing subjects.

The work is conducted within structured evaluation frameworks and professional guidelines. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality.

Role Details
  • Location: Seattle, WA

  • Compensation: $45-55

  • Employment classification: W-2

  • Schedule: 8AM - 5PM PT

  • Weekly commitment: M-F

Skills Required

  • Ability to learn customer policies, definitions, taxonomies, and evaluation rubrics quickly
  • Ability to evaluate user requests and AI model responses within full conversation context
  • Ability to distinguish closely related labels, severity levels, and policy boundaries
  • Ability to select defensible classifications in ambiguous cases
  • Ability to write concise, evidence-based rationales citing policy language and conversation details
  • Ability to identify policy gaps, contradictions, unclear definitions, and emerging edge cases
  • Ability to participate in calibration discussions and challenge interpretations respectfully
  • Ability to apply customer policy consistently without substituting personal beliefs
  • Accuracy and attention to detail during repeated evaluations
  • Ability to incorporate feedback quickly and apply clarified guidance
  • Ability to handle sensitive information and disturbing subject matter with maturity and sound judgment
  • Clear and precise written communication
  • Intellectual curiosity, self-direction, and comfort in a fast-changing environment
  • Ability to engage with sensitive material carefully, responsibly, and sustainably while maintaining consistent work quality
  • Professional experience evaluating or comparing ChatGPT, Claude, Gemini, or other language model outputs
  • Experience in AI evaluation, data annotation, RLHF, model quality, trust and safety, policy operations, or content moderation
  • Experience applying detailed rubrics, taxonomies, regulatory language, editorial standards, or quality frameworks
  • Familiarity with calibration sessions, inter-rater agreement, quality audits, or adjudication workflows
  • Experience writing policy guidance, decision trees, evaluation examples, or structured rationales
  • Comfort working with long conversations, incomplete context, and conflicting evidence
  • Familiarity with AI safety, responsible AI, and language-model risks

Handshake Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Handshake and has not been reviewed or approved by Handshake.

  • Leave & Time Off Breadth Time-off practices include flexible/unlimited PTO, companywide recharge weeks in summer and winter, plus additional volunteer and personal holiday time. Feedback suggests sabbaticals and coordinated breaks help people actually use rest time.
  • Parental & Family Support Parental leave is described as extended for primary and secondary caregivers, and family-oriented policies are highlighted. Fertility and family-support resources are referenced in public materials.
  • Healthcare Strength Core coverage spans medical, dental, and vision, with added mental-health resources and wellness programming. Feedback suggests these supports contribute meaningfully to overall wellbeing.

Handshake Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
700 Employees
Year Founded: 2014

What We Do

Handshake is the #1 place to launch a career with no connections, experience, or luck required. The platform connects up-and-coming talent with 650,000+ employers - from Fortune 500 companies like Google, Nike, and Target to thousands of public school districts, healthcare systems, and nonprofits. Earlier this year, we announced our $200M Series F funding round. This Series F fundraise and new valuation of $3.5B will fuel Handshake’s next phase of growth and propel our mission to help more people start, restart, and jumpstart their careers.

Why Work With Us

How someone builds their career is foundational. We believe in working with the higher education community to help students build meaningful careers. We are at the nexus of universities, students and employers— and we’re able to connect the very best pieces of each side. We’re proud to create a community where students can be more successful.

Gallery

Gallery

Similar Jobs

Cox Enterprises Logo Cox Enterprises

Sales Strategy & Enablement Director

Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Remote or Hybrid
United States
30000 Employees
135K-225K Annually
Hybrid
6 Locations
289097 Employees

Chewy Logo Chewy

Associate Director, Category Management

eCommerce • Healthtech • Pet • Retail • Pharmaceutical
Hybrid
Bellevue, WA, USA
17800 Employees
149K-245K Annually

Chewy Logo Chewy

Scientist

eCommerce • Healthtech • Pet • Retail • Pharmaceutical
Hybrid
Bellevue, WA, USA
17800 Employees
126K-190K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account