AI Engineer, Platforms

Reposted Yesterday
Be an Early Applicant
Singapore, SGP
In-Office
Junior
Artificial Intelligence • Healthtech • Information Technology • Biotech
The Role
Operate and optimize multi-cloud GPU inference platform for LLMs, build API services and CI/CD, manage AI clusters with IaC, improve observability, automate operations and build AI-assisted internal tools.
Summary Generated by Built In

AI Singapore (AISG) is a national AI programme launched by the National Research Foundation (NRF), Singapore, to build and anchor deep national capabilities in AI. AISG is supported through a government-wide partnership including the NRF, Ministry of Digital Development and Information (MDDI), Infocomm Media Development Authority (IMDA), Economic Development Board (EDB) and Enterprise Singapore (ESG). We bring together research institutions and the vibrant ecosystem of AI start-ups and companies to support impactful research, develop talent, and power Singapore's AI efforts.

This position will be hosted at the Nanyang Technological University (NTU) under VP (Artificial Intelligence & Digital Economy)’s office and we welcome you to join our community.

We're looking for an AI Engineer to join the Platform team within AI Products at AISG. In this role, you will be collaborating with different internal teams to design and implement optimized inference workflows and support the team to build customized LLM-based solutions.

Your work will directly contribute to the deployment, optimisation and management of large language models (LLMs) in production, retrieval-augmented generation (RAG) services, AI agent orchestration platforms and GPU-enabled AI infrastructure.

Responsibilities:

Platform operations and reliability

  • Own day-to-day operations of SEA-LION API Farm, our multi-cloud LLM inference platform — monitoring environment health, GPU capacity, performance, cost, and security posture.

  • Optimise LLM inference across various modalities to drive business value and support production goals.

  • Diagnose and troubleshoot performance and reliability issues on API Farm.

  • Build new API services such as batch API services, MCP services.

Infrastructure, CI/CD, and automation

  • Manage high performance AI clusters and storage systems using infrastructure-as-code (e.g. Terraform) across different cloud providers.

  • Develop and maintain CI/CD pipelines, container build/registry workflows, and deployment automation so teams can ship safely and frequently.

  • Strengthen observability across the stack including logs, metrics, traces, and dashboards, and reduce toil by automating repetitive operational tasks.

AI-assisted ops and continuous improvement

  • Use AI tools (e.g. Claude, Copilot, Cursor) appropriately in your daily work responsibilities.

  • Build internal tools leveraging AI to reduce manual effort in day-to-day operations.

Requirements:

You should be a hands-on engineer who is comfortable operating cloud and GPU infrastructure end-to-end, who understands how to deploy and run large language models reliably in production, and who actively uses AI tools to make platform work faster and more reliable.

  • A degree in Computer Science, Information Technology, or equivalent.

  • 1–3 years of DevOps, SRE, or platform engineering experience, with a track record of operating production systems at scale.

  • Hands-on experience operating workloads on different cloud providers including IaC (e.g. Terraform), containers and orchestration (e.g. Docker, Kubernetes), and managed services for compute, storage, and networking.

  • Strong knowledge on Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers).

  • Hands-on experience deploying and serving LLMs in production — model serving, GPU scheduling, autoscaling, latency/throughput optimisation, and inference cost management.

  • REST API design, model context protocol (MCP), Internet authentication patterns (e.g. OAuth).

  • Strong fundamentals in CI/CD, observability (logs/metrics/traces), and incident response.

  • Demonstrated use of AI tools (e.g. Claude, Copilot, Cursor) in your day-to-day engineering — for code generation, review, debugging, and documentation — with a clear sense of where they help and where they don't.

  • Solid scripting/programming skills (e.g. Python, Bash) and comfortable reading other people's code across the stack.

  • Strong communication skills with the ability to explain technical concepts.

Good to Have:

  • Experience with multimodal AI models (e.g. vision language models, audio language models).

  • C/C++/Rust/Go or other relevant programming languages.

  • Contributions to open-source AI/ML projects.

We regret that only shortlisted candidates will be notified.

Hiring Institution: NTU

Skills Required

  • Degree in Computer Science, Information Technology, or equivalent
  • 1-3 years of DevOps, SRE, or platform engineering experience operating production systems at scale
  • Infrastructure-as-code experience (e.g., Terraform) across cloud providers
  • Containers and orchestration experience (Docker, Kubernetes)
  • Hands-on experience deploying and serving LLMs in production (model serving, GPU scheduling, autoscaling, latency/throughput optimisation)
  • Experience with inference frameworks and libraries (vLLM, SGLang, TensorRT-LLM, Transformers)
  • REST API design and model context protocol (MCP); Internet authentication patterns (e.g., OAuth)
  • Strong fundamentals in CI/CD, observability (logs/metrics/traces), and incident response
  • Demonstrated use of AI tools (e.g., Claude, Copilot, Cursor) in day-to-day engineering
  • Solid scripting/programming skills (Python, Bash) and ability to read code across the stack
  • Strong communication skills able to explain technical concepts
  • Experience with multimodal AI models (vision/audio language models)
  • Experience with C, C++, Rust, Go or similar languages
  • Contributions to open-source AI/ML projects
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Singapore
10 Employees
Year Founded: 2020

What We Do

The Lee Kong Chian School of Medicine (LKCMedicine) trains doctors with a focus on patient-centered care, integrating precision medicine, Artificial Intelligence (AI) in healthcare, and medical humanities into its undergraduate medical degree program.

Similar Jobs

In-Office or Remote
17 Locations
5454 Employees

UL Solutions Logo UL Solutions

Senior Learning & Development Specialist

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Hybrid
Singapore, SGP
15000 Employees

Airwallex Logo Airwallex

Senior Software Engineer

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office
Singapore, SGP
2300 Employees

Qualtrics Logo Qualtrics

Account Executive

Artificial Intelligence • HR Tech • Information Technology • Software • Business Intelligence
Hybrid
Singapore, SGP
5000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account