This role is for one of our clients
Compensation: $50 - $200 per hour
A leading AI research organization is seeking advanced LLM power users to evaluate how well AI systems handle personalized, real-world life tasks.
This role is for people who use AI tools heavily in their personal lives and can clearly judge whether an AI response is useful, personalized, realistic, and successful.
Requirements
Who We’re Looking For
Strong candidates have:
- Heavy personal usage of LLM products
- Experience using AI for multi-step tasks, planning, research, decision-making, or personal workflows
- Familiarity with tools such as ChatGPT, Claude, Gemini, Perplexity, Cursor, Windsurf, Codex, or other AI agents
- Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic
- Strong written judgment and attention to detail
Why This Work Matters
LLMs are quickly becoming personal assistants for everyday decisions, but truly useful AI needs to do more than produce generic advice. It needs to understand context, preferences, constraints, tradeoffs, and what success looks like in real life. Your evaluations will help improve how AI systems support people with practical, high-context tasks across food, health, productivity, careers, and learning. This work directly contributes to making AI assistants more personalized, trustworthy, and useful for real-world personal workflows.
Engagement Details
- Expected commitment: 20-40 hours/week
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Skills Required
- Heavy personal usage of LLM products
- Experience using AI for multi-step tasks, planning, research, decision-making, or personal workflows
- Familiarity with AI tools such as ChatGPT, Claude, Gemini, Perplexity, Cursor, Windsurf, Codex, or other AI agents
- Ability to evaluate whether AI outputs are useful, personalized, realistic, successful, good, bad, incomplete, unsafe, or unrealistic
- Strong written judgment and attention to detail
- Availability for 20–40 hours per week
What We Do
Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.








