This role is for one of our clients
Compensation: $150 per hour
We are hiring residency-trained physicians across specialties for non-clinical work developing and evaluating clinical AI systems. You will apply your clinical judgment to grading-criteria development, dialogue evaluation, and structured annotation work that determines how these systems are measured.
This is a non-clinical role — no direct patient care, and no responsibility for live diagnosis.
This is a shared expert pool. After onboarding you may be matched to any of several concurrent clinical workstreams based on your specialty, availability, and interest. You are not committing to a single project, and you may move between streams as priorities shift.
Requirements
What you may work on
Work varies by workstream and may include:
- Grading criteria development — taking a clinical question and breaking the ideal answer into discrete, checkable criteria, so a model response can be graded consistently rather than impressionistically.
- Clinical dialogue evaluation — reviewing multi-turn clinical conversations and judging them for accuracy, safety, completeness, appropriate hedging, and whether escalation advice was correct.
- Clinical reasoning annotation — recording how you would work through a case, including the differential you considered and rejected, not only the conclusion.
- Output review — flagging hallucinated findings, dangerous omissions, unsupported certainty, and advice that is technically correct but clinically unsafe.
- Guideline authoring — defining edge cases and standards of care for your specialty so annotation stays consistent across a large group of clinicians.
- Difficult-case writing — constructing clinical questions that probe the limits of current model reasoning.
Task length varies by stream, from roughly 45 minutes for a dialogue evaluation up to an hour or more for grading-criteria authoring. You will get a specific throughput target for whichever stream you are matched to.
Required qualifications
- MD or DO with a completed residency in any specialty
- Active, unrestricted medical license in your country of practice
- 2+ years post-residency clinical experience, practising or previously practising
- Comfort writing structured clinical rationale that a non-specialist reviewer can follow
- Written and spoken English fluency
- Minimum 20 hours per week, with the ability to concentrate hours when a stream is time-boxed
Preferred qualifications
- Board certification in your specialty
- U.S. licensure, and familiarity with U.S. standards of care and clinical guidelines
- Primary care, internal medicine, emergency medicine, or hospitalist background, where breadth of presentation matters most
- Prior clinical annotation, AI evaluation, medical education, or question-writing experience
- Grading-criteria design, resident assessment, or clinical guideline development experience
- Published research or sustained technical writing (please link a sample)
Why this work
Most clinical AI failures are not exotic — they are ordinary questions answered with misplaced confidence. Catching that requires someone who has actually carried clinical responsibility. The standard these systems get held to is written by physicians, and here you would be writing it.
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Skills Required
- MD or DO degree
- Completed residency in any specialty
- Active, unrestricted medical license in the country of practice
- At least 2 years of post-residency clinical experience
- Ability to write structured clinical rationale understandable to non-specialists
- Written and spoken English fluency
- Availability for a minimum of 20 hours per week
- Board certification in the relevant specialty
- U.S. medical licensure and familiarity with U.S. standards of care and clinical guidelines
- Primary care, internal medicine, emergency medicine, or hospitalist background
- Clinical annotation, AI evaluation, medical education, or question-writing experience
- Grading-criteria design, resident assessment, or clinical guideline development experience
- Published research or sustained technical writing experience with a sample
What We Do
Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

.png)






