As an MLOps platform engineer, you will play a critical role by enabling our team of ML engineers to develop, train and deploy ML projects efficiently using the latest technologies in container orchestration, cloud services, CI/CD pipelines for data collection, model training and monitoring in production
Building and maintaining a scalable infrastructure to execute ML projects, while committed to results and user value
Develop, design, maintain and manage container orchestration using Kubernetes
Design and execute strategies for GPU optimization, prediction servers, data and training pipelines while ensuring efficient use
Design and build inference platforms while ensuring reliability and high performance
Provision and monitor infrastructure resources
Build and maintain ML workflows and pipelines
Deploy and maintain monitoring services for observability
Ensure compliance with security best practices
Manage and expand LLM serving clusters using stacks like vLLM
Requirements
Qualification
Bachelor's degree in Computer Science, engineering or related field
3+ years building core infrastructure for ML projects
Demonstrated background in DevOps, Platform Engineering, SRE, cloud-based infrastructure, or managing production operations
Experience supporting Generative AI, LLM, production-level AI/ML, or platforms focused on data-intensive workloads
Deep understanding of the AI application lifecycle, including MLOps, LLMOps, model monitoring, and deployment strategies
Hands-on experience deploying and providing support for AI services, inference endpoints, and APIs
Experience in managing, designing, implementing and maintaining robust ML infrastructure to support development and inference workloads, ML workflows, training pipelines and versioning
Experience building and scaling machine learning infrastructure
Experience with AWS cloud services
Experience with Kubernetes to deploy and manage containerized applications with high availability and performance
Experience in running and scaling inference clusters
Experience with TerraGrunt or TerraForm, IaC and CI/CD practices
Comfortable taking over legacy projects for operation and maintenance
Proficiency in programming Python
Excellent problem-solving skills and ability to work in a dynamic environment
Effective communication skills to collaborate with technical and nontechnical members
Master’s degree in Computer Science, engineering or related field
Production experience operating LLM inference servers such as vLLM (or equivalent serving stacks)
Experience with LLM observability, including the detection of hallucinations, toxicity, and model drift, alongside implementing tracing through OpenTelemetry protocols
Experience with RayServe
Proficiency on KubeFlow and MLFlow for workflows and pipelines
Experience in designing, developing and operating large-scale AI/ML systems
Certifications in AWS(MLS-C01), Kubernetes(CKA) or relevant technologies
Experience with additional cloud services
Contributions to open-source projects
Experience in working to improve model performance, including AI/ML model refinement and fine-tuning
Knowledge of data security standards such as handling personal information, financial/accounting data, PCI DSS, etc., and experience in designing, developing, and operating systems by these requirements.
Skills Required
- Bachelor's degree in Computer Science, engineering, or a related field
- 3+ years building core infrastructure for machine learning projects
- Experience in DevOps, platform engineering, SRE, cloud infrastructure, or production operations
- Experience supporting Generative AI, LLM, production-level AI/ML, or data-intensive platforms
- Understanding of MLOps, LLMOps, model monitoring, and deployment strategies
- Hands-on experience deploying and supporting AI services, inference endpoints, and APIs
- Experience designing and maintaining ML infrastructure, workflows, training pipelines, and versioning systems
- Experience building and scaling machine learning infrastructure
- Experience with AWS cloud services
- Experience deploying and managing highly available containerized applications with Kubernetes
- Experience running and scaling inference clusters
- Experience with Terraform or Terragrunt, Infrastructure as Code, and CI/CD practices
- Proficiency in Python programming
- Strong problem-solving skills
- Effective communication skills with technical and nontechnical stakeholders
- Master's degree in Computer Science, engineering, or a related field
- Production experience operating LLM inference servers such as vLLM or equivalent
- Experience with LLM observability, hallucination and toxicity detection, model drift, and OpenTelemetry tracing
- Experience with Ray Serve
- Proficiency with Kubeflow and MLflow
- Experience designing, developing, and operating large-scale AI/ML systems
- AWS MLS-C01, Kubernetes CKA, or relevant certifications
- Experience with additional cloud services
- Open-source project contributions
- Experience improving model performance through refinement and fine-tuning
- Knowledge of data security standards including personal information handling, financial data, and PCI DSS
What We Do
Money Forward, Inc. is a Japanese SaaS financial technology company that develops personal finance management tools and cloud-based services for individuals, businesses, accounting professionals, and financial institutions. Its platform supports household finance, accounting, tax, expense management, payments, and other back-office operations. The company serves consumer and corporate markets through integrated digital products and aims to operate as a comprehensive technology-driven financial platform.



.png)





