Senior Data Scientist ( Kernel Optimisation & Inference Engineer)

Posted Yesterday
Be an Early Applicant
Bengaluru, Karnataka, IND
In-Office
Senior level
Healthtech
The Role
Optimize GPU training and inference performance for medical language models. Responsibilities include profiling workloads, writing CUDA kernels, optimizing MoE communication and GEMMs, improving vLLM-class serving latency and throughput, implementing continuous batching and caching, validating AWQ/GPTQ/FP8 quantization, and enabling on-device inference for clinic hardware. The role requires hands-on GPU performance engineering, measurement-driven optimization, and correctness verification.
Summary Generated by Built In
Kernel Optimisation & Inference Engineer
Bengaluru  ·  Full-time  ·  2–5 yrs
Somewhere between the model and the silicon, 10–20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.

About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.

The role:
You'll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantised 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.

What you'll do
  • Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
  • Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.
  • Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
  • Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
  • Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.
  • Make on-device inference real for the hardware Indian clinics actually have.
What we look for
  • 2–5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.
  • Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain *why* FlashAttention is fast).
  • Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.
  • Measurement discipline: you profile before optimising and verify correctness after.
Bonus
  • fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.
  • Quantisation research or on-device/mobile inference experience.
  • Open-source kernels or serving contributions.
Why this is a rare gig
  • Open source, with your name on it: weights and technical reports ship publicly.
  • India-scale mission: models for a billion people in their own languages.
  • Compute that’s rare to fine: dedicated multi-node H200 training under a national grant.
  • Small senior team: you work with the people who own the recipe.
  • A live deployment path: Government institutes, EkaCare's doctors and patients use what you ship.

Full-Time Employee Benefits
  • Medical Insurance & Accidental Insurance
  • Maternity & Paternity Benefits
  • PF, Gratuity, & Leave Encashment
  • Salary Advance Policy

Skills Required

  • 2-5 years of experience in GPU performance work
  • Working fluency in CUDA
  • Strong memory-hierarchy reasoning, including coalescing, occupancy, and SRAM tiling
  • Hands-on experience with a modern model-serving stack such as vLLM, TensorRT-LLM, or SGLang
  • Experience profiling real workloads and shipping measurable optimizations
  • Ability to profile before optimizing and verify correctness afterward
  • FP8 experience on H100 or H200-class hardware
  • Experience with torch.compile or TorchInductor internals
  • Quantization research experience
  • On-device or mobile inference experience
  • Open-source kernel or model-serving contributions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Bengaluru, Karnataka
170 Employees
Year Founded: 2020

What We Do

A digitally enabled and connected healthcare ecosystem for better health management. - Manage Your Health Records - Monitor Your Health Vitals - Easy To Use - Private And Secured - Govt. of India Approved #prioritizehealth

Similar Jobs

Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
205000 Employees

Wells Fargo Logo Wells Fargo

Operations Specialist

Fintech • Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
205000 Employees

Wells Fargo Logo Wells Fargo

Operations Specialist

Fintech • Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
205000 Employees

Atlassian Logo Atlassian

Senior Engineering Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Bengaluru, Bengaluru Urban, Karnataka, IND
11000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
Vitalize Thumbnail
Artificial Intelligence • Healthtech • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account