Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Reposted One Month Ago
Be an Early Applicant
Menlo Park, CA, USA
In-Office
Senior level
Artificial Intelligence • Healthtech
The Role
Lead end-to-end post-training for LLMs using reinforcement learning and on-policy distillation to improve clinical reasoning, safety, and alignment. Design RL/OPD methods, build reward models and verifiers, create healthcare conversational environments with synthetic data, automate post-training loops, run experiments, and collaborate with research, engineering, and clinical teams to deploy models at scale.
Summary Generated by Built In
Role Mission

As HAI's LLM Post-Training Applied Scientist, you will own the reinforcement learning and on-policy distillation pipeline that transforms raw model capability into reliable, safe clinical behavior. Your post-training methods will directly determine how our AI agents reason through complex clinical scenarios, handle safety-critical decisions, and ultimately impact millions of patient interactions. This role exists because post-training is where capability becomes trustworthiness—and in healthcare, that's everything.

What You Will Accomplish

Own your first major outcome: By day 90, you will have shipped a post-training improvement that meaningfully advances model performance on a critical clinical capability (clinical reasoning, safety alignment, or task completion), evaluated the gains rigorously, and contributed that learning to our post-training roadmap.

Drive lasting impact: At 12 months, you will have designed and shipped multiple post-training methods that measurably improve our models' clinical safety and reasoning, built reusable infrastructure (reward models, verifiers, evaluation frameworks) that accelerate future post-training work, published your research or contributed to HAI's intellectual property, and directly shaped how our deployed models behave in production healthcare environments.

The Team

You'll work alongside ML researchers, engineers, clinicians, and safety experts who are obsessed with building trustworthy AI. This is a highly technical team that values rigor, collaboration across disciplines, and solving the hardest problems in AI safety and alignment. You'll have direct influence on model architecture and training decisions.

What You'll Do
  • Design and implement RL and OPD post-training methods including RLHF, RLVR, on-policy distillation, and novel approaches tailored to healthcare AI—selecting the right methods for different clinical reasoning and safety challenges

  • Build and evaluate reward models, verifiers, and LLM-as-judge pipelines that provide reliable training signals for post-training, ensuring they capture what truly matters in clinical contexts (accuracy, safety, patient experience)

  • Develop conversational AI environments and simulations for healthcare RL training—creating synthetic clinical scenarios and datasets that enable safe, scalable post-training without relying solely on human feedback

  • Automate post-training research loops using agents and tooling to systematically explore hyperparameters, methods, and data strategies—turning post-training into a scientific, reproducible process

  • Run rigorous experiments and analysis to understand what drives post-training gains, isolate the contributions of different components, and build intuition about what works in healthcare contexts

  • Collaborate with research, engineering, and clinical teams to translate clinical requirements into post-training objectives, validate improvements against real-world metrics, and scale successful methods to production

Location Requirement

We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Menlo Park office five days a week, unless otherwise specified.

Required Qualifications

  • Master's degree in Computer Science, Machine Learning, or a related field

  • 5+ years of professional experience in NLP, LLM training, or reinforcement learning

  • 2+ years of hands-on experience with RL for LLM post-training

  • Proficiency in Python and PyTorch for large-scale training

  • Demonstrated experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods

  • Experience training or fine-tuning models at scale (50B+ parameters)

Preferred Qualifications

  • Publications at top-tier ML venues (NeurIPS, ICML, ICLR, ACL, EMNLP)

  • Healthcare or regulated domain experience

  • Experience with distributed training frameworks (FSDP, DeepSpeed, vLLM)

  • Familiarity with safety alignment and interpretability research

Our comprehensive compensation package is designed to reward your expertise and includes both a competitive base salary and valuable stock options. Individual offers are determined based on a variety of factors, including your professional experience, core competencies, and geographic location.

Why Join Hippocratic AI

Reinvent healthcare with AI that puts safety first. We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact [email protected].

Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information during the hiring process.

Skills Required

  • MS or PhD in Computer Science or relevant field
  • 5+ years experience in NLP, LLM training, or reinforcement learning
  • 2+ years experience in RL for LLM post-training
  • Experience with large-scale (50B+ parameter and multi-node) LLM training
  • Strong Python coding skills
  • Strong PyTorch coding skills
  • Experience with RLHF, RLVR, LLM-as-judge or similar LLM post-training methods
  • On-site presence in Palo Alto office five days a week
  • Publications at top venues (NeurIPS, ICML, ICLR, ACL, EMNLP)
  • Healthcare domain experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, California
97 Employees
Year Founded: 2023

What We Do

Hippocratic AI’s mission is to develop the first safety-focused Large Language Model (LLM) for healthcare. The company believes that a safe LLM can dramatically improve healthcare accessibility and health outcomes in the world by bringing deep healthcare expertise to every human. No other technology has the potential to have this level of global impact on health. The company was co-founded by CEO Munjal Shah, alongside a group of physicians, hospital administrators, healthcare professionals, and artificial intelligence researchers from El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Microsoft, Meta and NVIDIA. Hippocratic AI has received a total of $137 million in funding and is backed by leading investors, including General Catalyst, Andreessen Horowitz, Premji Invest, SV Angel, NVentures (Nvidia Venture Capital), and Greycroft. For more information on Hippocratic AI: www.HippocraticAI.com.

Similar Jobs

Rapid7 Logo Rapid7

Senior Director, Customer Innovation

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
211K-285K Annually

Rapid7 Logo Rapid7

Vector Command Specialist

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
89K-121K Annually

Wipfli Logo Wipfli

Consultant

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
2900 Employees
117K-158K Annually

CrowdStrike Logo CrowdStrike

Sr. Security Researcher (Remote).

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
15 Locations
11000 Employees
85K-120K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account