MLOps / Serving Engineer

Posted 7 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Big Data • Cloud • Analytics • Consulting
The Role
Design and operate AWS production-serving infrastructure for fine-tuned LLMs. Deploy models using vLLM, TensorRT-LLM, or Triton; optimize inference and GPU utilization; implement shadow, canary, A/B, staged rollout, and automated rollback workflows; and build monitoring for latency, throughput, quality, and cost. The role also covers auto-scaling, high availability, Kubernetes-based ML workloads, continuous batching, quantization, and KV-cache management.
Summary Generated by Built In
Job Title: MLOps / Serving Engineer
Experience : 5+ years
Location : Hyderabad OR Pune
Notice Period: 0-30 days
Work mode - Hybrid
 
We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic.
Key Responsibilities:
  • Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances
  • Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic
  • Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation
  • Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization
  • Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request
  • Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics

Requirements
  • 5+ years MLOps or ML infrastructure engineering
  • Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server
  • Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling
  • Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback
  • Experience onto Continuous batching, INT8 quantization, KV-cache management
  • Expertise on Docker, Kubernetes (EKS) for ML workloads
  • Worked on CloudWatch, Prometheus, Grafana



Benefits
  • Comprehensive Medical Coverage:  
    Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:  
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits: 
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options: 
    Enjoy hybrid work arrangements & flexible working hours
  • Generous Leave Policy:  
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces: 
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Skills Required

  • 5+ years of experience in MLOps or ML infrastructure engineering
  • Hands-on experience with vLLM, TensorRT-LLM, or Triton Inference Server
  • Deep familiarity with AWS g5, p4de, and p5 instance families and EC2 Auto Scaling
  • Experience with shadow-mode deployments, canary releases, A/B traffic routing, and automated rollback
  • Experience with continuous batching, INT8 quantization, and KV-cache management
  • Expertise with Docker and Kubernetes, including EKS for ML workloads
  • Experience with CloudWatch, Prometheus, and Grafana
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
414 Employees
Year Founded: 2018

What We Do

DATAECONOMY is a global, cloud-first data and AI consultancy delivering enterprise-grade solutions through an innovative intellectual-property suite. Its work spans data and BI platform modernization, self-service AI, data mesh and fabric, master data management, governance, cloud enablement, digital engineering, knowledge graphs, and machine lakes supporting cybersecurity and financial-crime use cases for enterprise clients.

Similar Jobs

Optum Logo Optum

Full-stack Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Data Analyst

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Technical Product Manager

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Quality Engineer I - ACCELQ Selenium

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account