The Role
Design and run large-scale CPT and pre-training experiments for language models, including data-mix and curriculum ablations, proxy-scale testing, evaluation hooks, tokenization, long-context, and MoE experiments. Debug distributed training issues such as loss spikes, precision errors, dataloader stalls, and checkpoint failures. The role requires personally running pre-training or CPT at billion-parameter scale and applying rigorous empirical methods to optimize model recipes.
Summary Generated by Built In
ML Research Engineer; Pre-training (LLMs)
Bengaluru · Full-time · Experience: 2–4 yrs
About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
The role
The recipe is the game. You'll work at the heart of the 30B CPT: what goes into the ~500B-token mix, in what order, at what scale, proven cheaply at proxy scale, then spent confidently on the big run.
What you'll do
- Design and run CPT/mid-training ablation ladders at proxy scale, the experimental backbone that decides the real run.
- Own data-mix and curriculum empirics: medical vs general, Indic vs English, replay ratios (~70% design point), annealing schedules.
- Debug training at scale: loss spikes, precision issues, dataloader stalls, checkpoint pathologies.
- Build per-stage eval hooks so every CPT phase has a scoreboard, not a vibe.
- Run tokeniser, long-context and MoE-health experiments (routing balance, expert utilisation).
What we look for
- 2–4 years in ML with pretraining or CPT you personally ran at ≥1B scale (ideally ≥7B, 100B+ tokens) — the recipe was yours to break and fix.
- Fluency with Megatron/NeMo/TorchTitan-class trainers and distributed fundamentals (TP/PP/DP, mixed precision).
- Empirical rigour: you design ablations that answer questions.
- You read papers fast and implement faster.
Bonus
- MoE training exposure; scaling-laws mindset.
- Indic-language or domain-specific (medical/legal/code) pretraining.
- Kernels curiosity — you've opened a profiler and enjoyed it.
Why this is a rare gig
- Open source, with your name on it: weights and technical reports ship publicly.
- India-scale mission: models for a billion people in their own languages.
- Compute that’s rare to fine: dedicated multi-node H200 training under a national grant.
- Small senior team: you work with the people who own the recipe.
- A live deployment path: Government institutes, EkaCare's doctors and patients use what you ship.
Skills Required
- 2-4 years of machine learning experience
- Personally ran pre-training or continued pre-training at 1B+ parameter scale, ideally 7B+ with 100B+ tokens
- Fluency with Megatron, NeMo, TorchTitan, or comparable large-scale training frameworks
- Understanding of distributed training fundamentals, including tensor parallelism, pipeline parallelism, data parallelism, and mixed precision
- Ability to design rigorous empirical ablation experiments
- Ability to rapidly read research papers and implement methods
- Mixture-of-Experts training experience
- Scaling-laws experience or mindset
- Indic-language or domain-specific pre-training experience
- Experience using profilers and investigating kernel performance
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
A digitally enabled and connected healthcare ecosystem for better health management. - Manage Your Health Records - Monitor Your Health Vitals - Easy To Use - Private And Secured - Govt. of India Approved #prioritizehealth
.png)






