AI/ML MLOps Engineer

Posted 6 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Big Data • Cloud • Analytics • Consulting
The Role
Build and manage the full lifecycle for self-hosted LLMs, including training data pipelines, SFT and DPO fine-tuning, evaluation, benchmarking, quantization, deployment, monitoring, and continuous improvement. Deploy models on AWS GPU infrastructure and SageMaker, automate ML CI/CD, and manage Docker and Kubernetes workloads on EKS. Implement A/B, canary, and shadow deployments with automated promotion and rollback based on performance and operational metrics. Collaborate with data science, ML engineering, and DevOps teams.
Summary Generated by Built In

Job Title: AI/ML MLOps Engineer – LLM Fine-Tuning & Deployment

Experience: 5–8 Years

Location: Hyderabad
Employment Type: Full-Time, Hybrid
We are looking for an experienced AI/ML MLOps Engineer with strong hands-on expertise in LLM fine-tuning, model deployment, AWS GPU infrastructure, and MLOps. The role involves fine-tuning and deploying self-hosted Large Language Models (LLMs), building training and evaluation pipelines, and implementing reliable production deployment and monitoring practices.The ideal candidate should have practical experience working across the complete ML lifecycle — data preparation, model fine-tuning, evaluation, deployment, monitoring, and continuous improvement.

 

Key Responsibilities


  • Fine-tune Large Language Models using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO).
  • Develop and maintain training data pipelines, including data transformation, formatting, deduplication, filtering, and quality validation.
  • Work extensively with the Hugging Face ecosystem, including Transformers, Datasets, and PEFT.
  • Build and automate model evaluation and benchmarking frameworks to assess model quality and performance.
  • Deploy and serve LLM models using AWS GPU/EC2 infrastructure and Amazon SageMaker.
  • Optimize models for production through model quantization, inference optimization, and resource utilization.
  • Build robust MLOps and ML CI/CD pipelines covering model training, evaluation, packaging, deployment, and monitoring.
  • Implement A/B testing, Canary, and Shadow-mode deployments for safely introducing new model versions into production.
  • Develop mechanisms for automated model promotion and rollback based on predefined performance and operational metrics.
  • Implement production monitoring for model performance, latency, throughput, errors, GPU utilization, and resource consumption.
  • Containerize ML workloads using Docker and deploy/manage them using Kubernetes/Amazon EKS.
  • Collaborate with Data Scientists, ML Engineers, DevOps teams, and other stakeholders to build scalable and reliable AI/ML solutions.


Requirements

  • Strong programming experience in Python.
  • Hands-on experience with LLM fine-tuning, particularly SFT and DPO.
  • Strong knowledge of Hugging Face Transformers, Datasets, and PEFT.
  • Experience working with AWS GPU/EC2 and SageMaker for ML workloads.
  • Strong understanding of MLOps, ML CI/CD, and model lifecycle management.
  • Experience with LLM model serving and production deployment.
  • Experience building training data preparation and processing pipelines.
  • Knowledge of model evaluation, benchmarking, and performance optimization.
  • Hands-on experience with model quantization.
  • Experience implementing A/B, Canary, and Shadow-mode deployments


Benefits
  • Comprehensive Medical Coverage:  
    Health insurance of INR 7.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:  
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits: 
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options: 
    Enjoy hybrid work arrangements & flexible working hours
  • Generous Leave Policy:  
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces: 
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.


Skills Required

  • Strong programming experience in Python
  • Hands-on experience with LLM fine-tuning, particularly Supervised Fine-Tuning and Direct Preference Optimization
  • Strong knowledge of Hugging Face Transformers, Datasets, and PEFT
  • Experience with AWS GPU or EC2 infrastructure and Amazon SageMaker for ML workloads
  • Strong understanding of MLOps, ML CI/CD, and model lifecycle management
  • Experience with LLM model serving and production deployment
  • Experience building training data preparation and processing pipelines
  • Knowledge of model evaluation, benchmarking, and performance optimization
  • Hands-on experience with model quantization
  • Experience implementing A/B, canary, and shadow-mode deployments
  • Five to eight years of experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
414 Employees
Year Founded: 2018

What We Do

DATAECONOMY is a global, cloud-first data and AI consultancy delivering enterprise-grade solutions through an innovative intellectual-property suite. Its work spans data and BI platform modernization, self-service AI, data mesh and fabric, master data management, governance, cloud enablement, digital engineering, knowledge graphs, and machine lakes supporting cybersecurity and financial-crime use cases for enterprise clients.

Similar Jobs

Wells Fargo Logo Wells Fargo

Full-stack Engineer

Fintech • Financial Services
Hybrid
Hyderabad, Telangana, IND
205000 Employees

Wells Fargo Logo Wells Fargo

Operations Associate

Fintech • Financial Services
Hybrid
Hyderabad, Telangana, IND
205000 Employees
Hybrid
Hyderabad, Telangana, IND
205000 Employees
Hybrid
Hyderabad, Telangana, IND
205000 Employees

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account