Staff/Sr. ML Infrastructure / Platform Engineer

Posted 7 Days Ago
Be an Early Applicant
Hiring Remotely in Taipei, TWN
Remote
Senior level
Big Data • Cloud • Security • Software • Cybersecurity
The Role
Design, build, and operate a production GPU-accelerated LLM serving platform. Responsibilities include managing Kubernetes clusters and NVIDIA GPU nodes, optimizing multi-GPU inference, tuning autoscaling for cost and latency, maintaining Terraform and Helm infrastructure, and developing Prometheus and Grafana observability for utilization, cache performance, latency, costs, SLA violations, and OOM events. Bonus work includes LoRA/PEFT workflows, MLflow integration, adapter CI/CD, and advanced inference framework tuning.
Summary Generated by Built In

Join Trend ‧ Join New Generation

趨勢科技 - 全球雲端資安領航者 / 全亞洲最大軟體公司 / 企業版圖橫跨五大洲 / 趨勢全球研發基地在台灣 
===============================================================

About the Role 

We are building a production-grade, GPU-accelerated LLM serving platform that powers multiple AI products at enterprise scale. You will be responsible for designing, building, and operating the infrastructure that serves large language models — from raw Kubernetes cluster management to multi-GPU inference optimization and autoscaling. 

 

Required Qualifications 

Model Serving & Inference 

  • Operate multi-model LLM serving infrastructure 
  • Tune autoscaling policies to balance GPU cost and latency SLAs 

Kubernetes & GPU Infrastructure 

  • Operate production K8s clusters with NVIDIA GPU nodes 
  • Handle GPU node lifecycle: NVIDIA driver setup 

Infrastructure as Code 

  • Write and maintain Terraform/Terragrunt modules for AWS/GCP cloud 
  • Package platform components and model deployments as Helm charts 
  • Manage multi-environment configurations 

Observability & Performance 

  • Maintain monitoring stack: Prometheus, Grafana, 
  • Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token 
  • Set up alerting for SLA violations and OOM events 

 

Bonus Skills 

These are not required, but candidates with these skills will stand out. 

  • LoRA / PEFT fine-tuning workflows 
  • MLflow for experiment tracking, model registry, and automated adapter deployment 
  • Experience building LoRA adapter CI/CD pipelines (training → registry → serving) 
  • Experience with alternative inference frameworks such as SGLang or NVIDIA NIM, including deep Parameter Tuning for Continuous Batching, KV Cache management, and Speculative Decoding. 

===============================================================
連結智慧 守護世界 --- Connected Intelligence for Securing a Connected World

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tokyo
7,000 Employees

What We Do

We’re a global cybersecurity leader, helping to make the world safe for exchanging digital information. Fueled by decades of security expertise, global threat research, and continuous innovation, our cybersecurity platform protects hundreds of thousands of organizations and millions of individuals across clouds, networks, devices, and endpoints. As a leader in cloud and enterprise cybersecurity, our platform delivers a powerful range of advanced threat defense techniques optimized for environments like AWS, Microsoft, and Google, and central visibility for better, faster detection and response. Our global threat research team delivers unparalleled intelligence and insights that power our cybersecurity platform and help protect organizations around the world from 100s of millions of threats daily. We have 7,000 employees across 65 countries, singularly focused on security and passionate about making the world a safer and better place. We enable organizations to simplify and secure their connected world. Trend Micro’s “Trenders” are passionate about doing the right thing to make the world a safer and better place.

Similar Jobs

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Thermal Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
Remote
Taipei, TWN
85422 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Design Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
Remote or Hybrid
Taipei, TWN
85422 Employees

UL Solutions Logo UL Solutions

Project Engineer

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
台灣
15000 Employees

The Aerospace Corporation Logo The Aerospace Corporation

Advanced Mech Design Engr

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
Taiwan
4600 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account