The Role
Develop and post-train large language models for safe, instruction-following clinical applications across Indian languages. Responsibilities include designing SFT and preference-optimization datasets, training reward models, building physician-guided reinforcement-learning loops, creating healthcare tool-use data, and conducting evaluations, error analysis, ablations, and red-teaming. Requires production experience with open-weight models, post-training frameworks, Python, distributed training, and strong empirical practices.
Summary Generated by Built In
ML Research Engineer; Post-training (LLMs)
Bengaluru · Full-time · Experience: 4–8 yrs
India's healthcare runs in twenty-two languages, on handwritten prescriptions and ten-minute consults, and the models that should serve it are trained on the English internet. We're fixing that, in the open.
About EkaCare
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
The role
Post-training is where a base model becomes a doctor's tool, and where most medical models quietly fail. You'll turn a strong pre-trained base into a model that follows instructions, knows what it doesn't know, stays safe in a clinical setting, and does it in a dozen Indian languages.
What you'll do
- Design SFT/IFT data mixes and chat templates for clinical tasks, and reformat the world's medical data to match.
- Run preference optimisation (DPO/ORPO/GRPO-class) and reward-model training; own the ablation grid.
- Build RL loops with verifiable medical rewards, with practising physicians in the loop. Real doctor-in-the-loop, not proxy labels.
- Create agentic and tool-use training data for healthcare workflows.
- Live in the eval → error-analysis → iterate loop; red-team your own model before the world does.
What we look for
- 2–4 years in ML with hands-on post-training of ≥7B open-weights models; you've shipped SFT plus at least one preference-optimisation method end to end, not a notebook demo.
- Fluency with the open post-training stack (TRL / NeMo RL / OpenRLHF; vLLM for rollouts).
- Strong empirical taste: ablation discipline, LLM-judge literacy, contamination paranoia.
- Solid engineering, Python, distributed-training basics, comfort in a fast codebase.
Bonus
- PPO/GRPO at scale; reward-hacking war stories.
- Multilingual or medical/clinical alignment work.
- Open-source contributions people actually use.
Skills Required
- 2-4 years of machine learning experience
- Hands-on post-training experience with open-weight models of at least 7B parameters
- End-to-end experience shipping supervised fine-tuning and at least one preference-optimization method
- Fluency with TRL, NeMo RL, or OpenRLHF
- Experience using vLLM for rollouts
- Strong Python engineering skills
- Distributed-training fundamentals
- Experience with PPO or GRPO at scale
- Multilingual or medical/clinical alignment experience
- Useful open-source contributions
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
A digitally enabled and connected healthcare ecosystem for better health management. - Manage Your Health Records - Monitor Your Health Vitals - Easy To Use - Private And Secured - Govt. of India Approved #prioritizehealth
.png)






