Senior DevOps/ MLOps Engineer

Posted 13 Days Ago
Be an Early Applicant
Hà Nội, VNM
In-Office
Senior level
Automotive • Retail • Transportation • Manufacturing
The Role
Design, deploy, and optimize scalable MLOps/model-serving infrastructure and GPU compute for low-latency Agentic AI and VoiceAI. Build EKS-based Kubernetes clusters, IaC (Terraform/CloudFormation), CI/CD pipelines, observability, and GPU performance tuning to productionize large AI models.
Summary Generated by Built In

VINFAST is a pioneering electric vehicle (EV) company committed to revolutionizing the automotive industry with sustainable and innovative mobility solutions. As a leading player in the EV market, VinFast is dedicated to delivering high-quality, cutting-edge electric vehicles that redefine the driving experience. Our team consists of passionate professionals driven by a shared vision of creating a greener and more sustainable future through innovation, technology, and excellence.  

The Speech & Language Processing (SLP) Center – Automotive AI Development Institute, is responsible for researching, developing, and deploying advanced Agentic AI and VoiceAI solutions, applied in vehicles and extended to other group-wide use cases such as robotics and AI assistants. 

We're looking for people with relevant experience, passion, and drive — ready to challenge themselves, keep learning, and thrive under high pressure to help build innovative products. 


Position Overview: This critical role serves as the backbone of infrastructure execution for Vingroup, directly accelerating the productionization of advanced AI capabilities across the global ecosystem. By architecting scalable MLOps pipelines and optimizing high-performance GPU infrastructure, this position ensures VinFast’s intelligent voice, Agentic AI, and robotics solutions run with maximum efficiency and near-zero latency. The role drives engineering excellence within the AI Model Deployment Department, translating heavy AI models into lean, resilient, and highly available production services. Ultimately, this position safeguards Vingroup’s technological velocity by scaling cutting-edge autonomous and smart solutions globally. 

In this role, you will be instrumental in AI Model Deployment Department, using your skills to build, scale, and optimize the core infrastructure that powers our next-generation Agentic AI and VoiceAI systems. As a Senior DevOps/MLOps Engineer, you will own the end-to-end deployment lifecycle, mastering model serving frameworks and advanced GPU compute orchestration. You will be responsible for transforming complex AI architectures into production-ready, ultra-low-latency services running on enterprise cloud and hybrid environments. 

You will collaborate with diverse teams, including AI Research Scientists, Technical Project Managers, Embedded Systems Developers, Cloud Architects, and Core Product Teams, to create cutting-edge solutions that will drive the future of transportation. 


 

Model Serving & GPU Optimization 

  • Design, deploy, and maintain high-performance model serving infrastructures utilizing advanced engines such as vLLM and Triton Inference Server to support Large Language Models (LLMs) and VoiceAI systems. 
  • Implement deep GPU optimization strategies (e.g., dynamic batching, quantization, tensor parallelism, and memory management) to maximize hardware utilization, minimize time-to-first-token (TTFT), and reduce VRAM overhead. 
  • Architect auto-scaling policies based on custom metrics (such as concurrent requests, GPU queue time, and VRAM limits) to handle highly fluctuating traffic patterns smoothly. 
  • Benchmark and profile model inference workloads, identifying infrastructure bottlenecks between compute, memory, and network layers. 

Cloud Infrastructure & Platform Engineering 

  • Build and orchestrate enterprise-grade Kubernetes clusters on AWS using EKS, ensuring robust network routing, service mesh integration, and strict security compliance. 
  • Develop and maintain Infrastructure as Code (IaC) using tools like Terraform or CloudFormation to ensure reproducible, predictable environments across Dev, Staging, and Production. 
  • Design CI/CD pipelines optimized for MLOps, automating the automated testing, packaging, and seamless canary/blue-green deployment of both application code and heavy AI weights. 
  • Establish comprehensive observability frameworks (Prometheus, Grafana, ELK stack, or equivalent) tracking system health, cluster performance, and specialized GPU/model telemetry. 



Requirements

Education & Background 

  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related technical discipline. 
  • AWS Certified DevOps Engineer or Kubernetes certifications (CKA/CKAD) are highly advantageous. 

Work Experience 

  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering. 
  • 2+ years of hands-on experience specifically in MLOps roles, deploying and maintaining large-scale AI/ML models in production. 
  • Proven track record of managing production-grade clusters on AWS (EKS) handling high-throughput, low-latency traffic. 

Technical Knowledge & Expertise 

  • In-depth expertise in AWS & EKS: Master of VPC networking, IAM policies, EKS node groups (including GPU-accelerated instances like p4/g5 families), and cluster security. 

  • Advanced Model Serving: Proficiency in configuring and tuning vLLM and Triton Inference Server for optimal concurrent execution. 

  • Deep GPU Infrastructure Mastery: Solid understanding of NVIDIA CUDA environments, driver management, multi-instance GPUs (MIG), and hardware-level performance tuning. 

  • CI/CD & Automation: Strong expertise in modern pipeline engines (GitLab CI, GitHub Actions, or ArgoCD) and container management. 

  • Scripting & Frameworks: Strong proficiency in Python and Bash for infrastructure automation, workflow orchestration, and tooling development. 



Benefits
  • Competitive salary
  • Premium healthcare package, including PVI insurance & annual health check-ups
  • 13th-month salary & performance bonuses to reward your contributions
  • Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool...
  • Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain
Work Location: Technopark Tower, Vinhomes Ocean Park, Gia Lam, Hanoi, Vietnam

With respect to all your personal data shared to VinFast in the application and the entire recruitment process of VinFast, by clicking “Apply”, submitting your resumé/CV and/or participating in VinFast's recruitment process, you agree that you have read VinFast's Personal Data Protection Policy ("Policy") posted at https://vinfastauto.com/vn_vi/dieu-khoan-phap-ly or https://vinfast.vn/privacy-policy/, you agree to the Policy and consent for VinFast to process your personal data in accordance with the Policy and the applicable regulations on personal data protection.

Skills Required

  • Bachelor's degree in Computer Science, Software Engineering, IT, or related field
  • 5+ years in DevOps, SRE, or Infrastructure Engineering
  • 2+ years hands-on MLOps experience deploying large-scale AI/ML models
  • Proven experience managing production-grade AWS EKS clusters
  • In-depth AWS & EKS expertise (VPC networking, IAM, GPU node groups)
  • Experience configuring and tuning vLLM and Triton Inference Server
  • Deep GPU infrastructure knowledge (NVIDIA CUDA, driver management, MIG)
  • Experience with IaC tools (Terraform or CloudFormation)
  • CI/CD and automation expertise (GitLab CI, GitHub Actions, or ArgoCD)
  • Observability tooling experience (Prometheus, Grafana, ELK or equivalent)
  • Strong scripting skills in Python and Bash
  • AWS Certified DevOps Engineer or Kubernetes certifications (CKA/CKAD)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Cát Hải
29,878 Employees
Year Founded: 2017

What We Do

VinFast is a Vietnamese multinational automotive manufacturer, established in 2017, that designs and manufactures electric vehicles (EVs), e-scooters, and e-buses. It is part of Vingroup, one of Vietnam's largest conglomerates.

Similar Jobs

UL Solutions Logo UL Solutions

Laboratory Engineer Associate

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
Việt Nam
15000 Employees
Remote or Hybrid
2 Locations
289097 Employees

UL Solutions Logo UL Solutions

Intern, Chemical Safety Testing

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
Việt Nam
15000 Employees

UL Solutions Logo UL Solutions

Intern, Textile Testing (Softline)

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
Việt Nam
15000 Employees

Similar Companies Hiring

Fortune Brands Innovations Thumbnail
Manufacturing
Deerfield, IL
10000 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account