AI ML Ops Engineer

Posted 4 Days Ago
Be an Early Applicant
Indore, Madhya Pradesh, IND
In-Office
Mid level
Big Data • Information Technology • Software • Consulting
The Role
Owns the deployment, monitoring, scaling, and reliability of machine learning models and production pipelines. Responsibilities include building ML CI/CD workflows, containerizing and orchestrating services with Docker and Kubernetes, managing cloud GPU infrastructure, tracking experiments and models, monitoring drift and system health, and optimizing inference performance using Triton, TensorRT, and ONNX Runtime. The role also supports LLM inference, infrastructure as code, observability, and bare-metal Linux troubleshooting.
Summary Generated by Built In
About Tarento:

Tarento is a fast-growing technology consulting company headquartered in Stockholm, with a strong presence in India and clients across the globe. We specialize in digital transformation, product engineering, and enterprise solutions, working across diverse industries including retail, manufacturing, and healthcare. Our teams combine Nordic values with Indian expertise to deliver innovative, scalable, and high-impact solutions.

We're proud to be recognized as a Great Place to Work, a testament to our inclusive culture, strong leadership, and commitment to employee well-being and growth. At Tarento, you’ll be part of a collaborative environment where ideas are valued, learning is continuous, and careers are built on passion and purpose.

Designation: Senior Software Engineer – Analytics
Location: Indore
Educational Qualifications: B.E/B.Tech
Exp: 3-5 Years
Mode of work: Hybrid

Role Overview
We are seeking an AI/ML Ops Engineer to own the deployment, monitoring, and scaling of ML models and pipelines in production. You will build the infrastructure and automation that lets models move reliably from training to production and stay healthy once there.

Key Responsibilities
  • CI/CD for ML: Build and maintain automated pipelines for model training, testing, and deployment.
  • Infrastructure: Containerize and orchestrate model-serving infrastructure (Docker, Kubernetes) at scale.
  • Monitoring: Set up monitoring for model performance, data/concept drift, latency, and system health.
  • Versioning & Reproducibility: Manage model/experiment versioning and reproducibility (MLflow, DVC, or similar).
  • Cloud & GPU Management: Manage scalable, cost-efficient GPU/cloud infrastructure for training and inference workloads.
Technical Requirements
  • Strong hands-on experience with Docker and Kubernetes
  • CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, or similar)
  • Experience with model-serving frameworks (Triton Inference Server, TorchServe, or similar)
  • Cloud platform experience (AWS/Azure/GCP), especially GPU infrastructure
  • Python for automation and tooling
  • Experience with experiment tracking and model registries (MLflow, DVC, Weights & Biases)
  • Monitoring/observability tooling (Prometheus, Grafana, or similar)
  • Production experience with Triton Inference Server: model repositories, ensembles, dynamic batching, instance groups
  • GPU operations on Kubernetes: NVIDIA GPU Operator, MIG/time-slicing, node pools, driver/CUDA version management
  • Hands-on with TensorRT / ONNX Runtime conversion and performance profiling (Nsight, perf_analyzer) gRPC and streaming service patterns; load testing tools (Locust, k6)
  • Strong Linux, networking and debugging fundamentals for bare-metal environments
Mandatory Skills
  • Docker, Kubernetes, CI/CD, Python, cloud infrastructure (AWS/Azure/GCP)
Good to Have
  • Experience deploying LLM/NLP inference services specifically (batching, quantization, low-latency serving)
  • Familiarity with vLLM, Ollama, or similar for local LLM hosting
  • Infrastructure-as-code experience (Terraform, Helm)
  • Experience with NVIDIA NVCF, NeMo or NIM deployments
  • Serving TTS/ASR models where time-to-first-byte matters
  • Log/trace stacks (Loki, OpenTelemetry, ELK) and on-call tooling

Skills Required

  • B.E. or B.Tech degree
  • 3-5 years of professional experience
  • Hands-on experience with Docker and Kubernetes
  • Experience with CI/CD tooling such as GitHub Actions, Jenkins, or GitLab CI
  • Experience with model-serving frameworks such as Triton Inference Server or TorchServe
  • Cloud platform experience with AWS, Azure, or GCP, especially GPU infrastructure
  • Python for automation and tooling
  • Experience with experiment tracking and model registries such as MLflow, DVC, or Weights & Biases
  • Experience with monitoring and observability tooling such as Prometheus or Grafana
  • Production experience with Triton Inference Server, including model repositories, ensembles, dynamic batching, and instance groups
  • GPU operations on Kubernetes, including NVIDIA GPU Operator, MIG or time-slicing, node pools, and driver or CUDA management
  • Hands-on experience with TensorRT or ONNX Runtime conversion and performance profiling
  • Experience with gRPC, streaming service patterns, and load-testing tools such as Locust or k6
  • Strong Linux, networking, and debugging fundamentals for bare-metal environments
  • Experience deploying LLM or NLP inference services
  • Familiarity with vLLM, Ollama, or similar local LLM hosting tools
  • Infrastructure-as-code experience with Terraform or Helm
  • Experience with NVIDIA NVCF, NeMo, or NIM deployments
  • Experience serving TTS or ASR models
  • Familiarity with Loki, OpenTelemetry, ELK, and on-call tooling
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
766 Employees
Year Founded: 2009

What We Do

Tarento is a Nordic-Indian IT services company that helps organizations navigate digitalization through enterprise applications, data and information management, mobile and custom solutions. Its offerings span digital, data, AI, cloud, consulting, software development, integration, testing, and application management. The company also builds scalable web and mobile apps, portals, platforms, and products for clients across industries, combining technology expertise with data-management capabilities.

Similar Jobs

Capco Logo Capco

Accounts BA

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Capco Logo Capco

BA - Advisory (Wealth) GCB4

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Capco Logo Capco

Product Manager

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

CSC Logo CSC

Accountant

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account