The Role
Build large-scale LLM training infrastructure, including corpus pipelines, distributed training workflows, evaluation tooling, experiment systems, and high-performance inference paths. The role involves training or fine-tuning billion-parameter models, operating multi-node H200 clusters, improving observability and recovery, and automating machine learning workflows. Experience with distributed computing, large-scale data engineering, and production ML systems is required.
Summary Generated by Built In
Senior ML Engineer, LLM Systems
Bengaluru · Full-time · mid-senior (3–6 yrs)
About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
The role
The versatile builder, you go where the run needs you; data, training, evals, serving, and you leave every system better-instrumented than you found it.
What you'll do
- Build corpus pipelines that feed GPUs without stalls: filtering, dedup, PII redaction, tokenisation, streaming dataloaders at ~500B-token scale.
- Run distributed training jobs on multi-node H200 clusters; make failures loud, recoveries fast, dashboards honest.
- Build the eval and experiment tooling researchers live in.
- Wire fast inference paths (vLLM-class) for model iterations.
- Automate whatever annoyed you last week.
What we look for
- 3–6 years of strong software engineering with real ML training exposure; you've fine-tuned or trained ≥1B-parameter models, not just called APIs.
- Distributed-computing fundamentals: Slurm or Kubernetes, NCCL basics, storage/throughput intuition.
- Data engineering at scale; Spark, Ray, or whatever gets the job done.
- A bias for ownership: you don't ask whose job it is.
Bonus
- Triton/CUDA curiosity; profiler literacy.
- Open-source projects or contributions.
- Healthcare or Indic-language data experience.
Why this is a rare gig
- Open source, with your name on it: weights and technical reports ship publicly.
- India-scale mission: models for a billion people in their own languages.
- Compute that’s rare to fine: dedicated multi-node H200 training under a national grant.
- Small senior team: you work with the people who own the recipe.
- A live deployment path: Government institutes, EkaCare's doctors and patients use what you ship.
Full-Time Employee Benefits
- Medical Insurance & Accidental Insurance
- Maternity & Paternity Benefits
- PF, Gratuity, & Leave Encashment
- Salary Advance Policy
Skills Required
- 3-6 years of strong software engineering experience
- Real machine learning training experience, including fine-tuning or training at least one 1B-parameter model
- Distributed-computing fundamentals, including Slurm or Kubernetes
- Understanding of NCCL, storage, and system throughput
- Large-scale data engineering experience using Spark, Ray, or comparable technologies
- Triton or CUDA experience
- Profiler literacy
- Open-source projects or contributions
- Healthcare or Indic-language data experience
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
A digitally enabled and connected healthcare ecosystem for better health management. - Manage Your Health Records - Monitor Your Health Vitals - Easy To Use - Private And Secured - Govt. of India Approved #prioritizehealth








