MLOps Platform Engineer (Chennai / Pune)

Posted 15 Hours Ago
Be an Early Applicant
3 Locations
In-Office or Remote
Mid level
Cloud • Fintech • Software
The Role
Build and operate scalable MLOps infrastructure for machine learning development, training, inference, and monitoring. Responsibilities include Kubernetes orchestration, AWS resource management, GPU optimization, prediction and inference platforms, ML workflows, CI/CD pipelines, observability, security compliance, and LLM serving clusters. The role collaborates with ML engineers and supports reliable, high-performance AI services and production operations.
Summary Generated by Built In
Overview
Money Forward is developing a variety of services for individuals and corporations to realize our vision, “Becoming the financial platform for all”. In addition, we are working to promote the effective use of data. To further address our customers' needs in the future, we are actively strengthening our development system using AI/ML technology for the main services of each department.
We are looking for a passionate Platform Engineer for MLOps who can work along with our ML platform team and collaborate with ML engineers to ensure deployment, scaling, and maintenance of ML pipelines and applications.
You will help manage cloud-based resources, containerized environments, and automated workflows for AI models. You will contribute to building a scalable AI/ML infrastructure while gaining exposure to the broader MLOps lifecycle and automation.

Attractive points
In this role, you will be at the forefront of the latest technologies in container orchestration, cloud services, and CI/CD pipelines to enable efficient development, training and deployment of ML models.

You will have the autonomy to design and implement optimization strategies, operate and maintain a scalable robust infrastructure tailored for ML projects, and empower ML engineers throughout the MLOps cycle.

Alongside our technical team of talented experienced ML engineers, you will also have the opportunity to contribute to the MLOps cycle, gaining valuable insights in a diverse and dynamic environment.
Responsibilities
  • As an MLOps platform engineer, you will play a critical role by enabling our team of ML engineers to develop, train and deploy ML projects efficiently using the latest technologies in container orchestration, cloud services, CI/CD pipelines for data collection, model training and monitoring in production

  • Building and maintaining a scalable infrastructure to execute ML projects, while committed to results and user value

  • Develop, design, maintain and manage container orchestration using Kubernetes

  • Design and execute strategies for GPU optimization, prediction servers, data and training pipelines while ensuring efficient use

  • Design and build inference platforms while ensuring reliability and high performance

  • Provision and monitor infrastructure resources 

  • Build and maintain ML workflows and pipelines

  • Deploy and maintain monitoring services for observability 

  • Ensure compliance with security best practices 

  • Manage and expand LLM serving clusters using stacks like vLLM



Requirements

Qualification

  • Bachelor's degree in Computer Science, engineering or related field

  • 3+ years building core infrastructure for ML projects

  • Demonstrated background in DevOps, Platform Engineering, SRE, cloud-based infrastructure, or managing production operations

  • Experience supporting Generative AI, LLM, production-level AI/ML, or platforms focused on data-intensive workloads

  • Deep understanding of the AI application lifecycle, including MLOps, LLMOps, model monitoring, and deployment strategies

  • Hands-on experience deploying and providing support for AI services, inference endpoints, and APIs

  • Experience in managing, designing, implementing and maintaining robust ML infrastructure to support development and inference workloads, ML workflows, training pipelines and versioning

  • Experience building and scaling machine learning infrastructure

  • Experience with AWS cloud services

  • Experience with Kubernetes to deploy and manage containerized applications with high availability and performance

  • Experience in running and scaling inference clusters

  • Experience with TerraGrunt or TerraForm, IaC and CI/CD practices

  • Comfortable taking over legacy projects for operation and maintenance

  • Proficiency in programming Python

  • Excellent problem-solving skills and ability to work in a dynamic environment

  • Effective communication skills to collaborate with technical and nontechnical members

Nice-to-have
  • Master’s degree in Computer Science, engineering or related field

  • Production experience operating LLM inference servers such as vLLM (or equivalent serving stacks)

  • Experience with LLM observability, including the detection of hallucinations, toxicity, and model drift, alongside implementing tracing through OpenTelemetry protocols

  • Experience with RayServe

  • Proficiency on KubeFlow and MLFlow for workflows and pipelines

  • Experience in designing, developing and operating large-scale AI/ML systems

  • Certifications in AWS(MLS-C01), Kubernetes(CKA) or relevant technologies

  • Experience with additional cloud services

  • Contributions to open-source projects

  • Experience in working to improve model performance, including AI/ML model refinement and fine-tuning

  • Knowledge of data security standards such as handling personal information, financial/accounting data, PCI DSS, etc., and experience in designing, developing, and operating systems by these requirements.


Skills Required

  • Bachelor's degree in Computer Science, engineering, or a related field
  • 3+ years building core infrastructure for machine learning projects
  • Experience in DevOps, platform engineering, SRE, cloud infrastructure, or production operations
  • Experience supporting Generative AI, LLM, production-level AI/ML, or data-intensive platforms
  • Understanding of MLOps, LLMOps, model monitoring, and deployment strategies
  • Hands-on experience deploying and supporting AI services, inference endpoints, and APIs
  • Experience designing and maintaining ML infrastructure, workflows, training pipelines, and versioning systems
  • Experience building and scaling machine learning infrastructure
  • Experience with AWS cloud services
  • Experience deploying and managing highly available containerized applications with Kubernetes
  • Experience running and scaling inference clusters
  • Experience with Terraform or Terragrunt, Infrastructure as Code, and CI/CD practices
  • Proficiency in Python programming
  • Strong problem-solving skills
  • Effective communication skills with technical and nontechnical stakeholders
  • Master's degree in Computer Science, engineering, or a related field
  • Production experience operating LLM inference servers such as vLLM or equivalent
  • Experience with LLM observability, hallucination and toxicity detection, model drift, and OpenTelemetry tracing
  • Experience with Ray Serve
  • Proficiency with Kubeflow and MLflow
  • Experience designing, developing, and operating large-scale AI/ML systems
  • AWS MLS-C01, Kubernetes CKA, or relevant certifications
  • Experience with additional cloud services
  • Open-source project contributions
  • Experience improving model performance through refinement and fine-tuning
  • Knowledge of data security standards including personal information handling, financial data, and PCI DSS
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
2,962 Employees
Year Founded: 2012

What We Do

Money Forward, Inc. is a Japanese SaaS financial technology company that develops personal finance management tools and cloud-based services for individuals, businesses, accounting professionals, and financial institutions. Its platform supports household finance, accounting, tax, expense management, payments, and other back-office operations. The company serves consumer and corporate markets through integrated digital products and aims to operate as a comprehensive technology-driven financial platform.

Similar Jobs

Remote or Hybrid
2 Locations
289097 Employees

Motive Logo Motive

Operations Manager

Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Easy Apply
Remote
India
4000 Employees

Capco Logo Capco

Liquidity Reporting

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Sailor Health Logo Sailor Health

Care Coordinator

Healthtech • Social Impact • Telehealth
In-Office or Remote
6 Locations
20 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account