AI Infrastructure Hardware Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office or Remote
Entry level
Information Technology • Software
The Role
Own quality control, hardware bring-up, validation, stress testing, and end-to-end testing of GPU/HPC servers across factory and lab environments. Diagnose hardware, firmware, driver, thermal, power, and interconnect issues; validate CUDA and NCCL performance under Linux; define test procedures and acceptance criteria; document results; and coordinate with manufacturing partners and suppliers in English and Mandarin.
Summary Generated by Built In

Job Description

Location: Hong Kong — Hybrid (factory floor, test lab and manufacturing sites; this is not a remote or desk-based role)
Start date: ASAP
Languages: Fluent English and fluent Mandarin are both mandatory
Industry: HPC / GPU Server Manufacturing / Sovereign Cloud Infrastructure

Pragmatike is recruiting on behalf of a European deep-tech company building a sovereign, energy-efficient alternative to traditional cloud providers. Unlike most cloud players, our client designs and manufactures its own HPC infrastructure end to end — chassis, boards and containerized units — and deploys it as the hardware layer of a vertically integrated AI cloud. The company is growing fast and is expanding its manufacturing and testing capability in Hong Kong.

Professional Background:

We're looking for a Server Engineer with hands-on GPU/HPC hardware experience who has tested and validated servers, not just operated them. You should be comfortable moving between the factory floor and the test lab, with a strong Linux background and real familiarity with firmware and GPU software stacks (CUDA, NCCL). You take ownership of quality: if a server ships, it is because you signed off on it.
Startup or hyper-growth experience is a strong plus. Autonomy, rigor and problem-solving are essential.

Your Responsibilities:

  • Own quality control across the client's manufacturing sites, on the factory floor.

  • Run the hardware test lab: board and system bring-up, validation and stress testing of new server builds.

  • Oversee the assembly and end-to-end testing of GPU servers before they are deployed into the client's sovereign cloud.

  • Define and maintain test procedures, acceptance criteria and QC checkpoints for chassis, boards and full systems.

  • Diagnose hardware, firmware and driver-level issues (BIOS/BMC, GPU firmware, PCIe, thermals, power) and drive root-cause analysis with the design and manufacturing teams.

  • Validate GPU compute and interconnect performance under Linux using CUDA, NCCL and related benchmarking tools.

  • Document failures, yields and test results, and feed findings back into hardware design and production processes.

  • Coordinate day to day with manufacturing partners and local suppliers in both English and Mandarin.

What You Bring:

  • Proven hands-on experience testing and validating GPU/HPC servers (bring-up, burn-in, validation), not only using them.

  • Strong Linux skills at the system and hardware level.

  • Firmware knowledge: BIOS/UEFI, BMC/IPMI, GPU and NIC firmware, flashing and troubleshooting.

  • Working knowledge of the NVIDIA GPU stack: drivers, CUDA, NCCL, and how to benchmark and diagnose multi-GPU systems.

  • Solid understanding of server hardware: motherboards, PCIe topology, power, cooling, storage and networking components.

  • Experience with QC or manufacturing test processes in a hardware production environment.

  • Fluent English and fluent Mandarin (both non-negotiable).

  • Willingness to work on-site in Hong Kong, on the factory floor and in the lab.

Nice To Have:

  • Experience with containerized or micro-datacenter deployments (power, thermal and environmental constraints).

  • Familiarity with high-speed interconnects (InfiniBand, RoCE, NVLink) and their validation.

  • Scripting skills (Bash/Python) to automate test sequences and reporting.

  • Exposure to hardware supply chains or contract manufacturers in Hong Kong / Southern China.

  • Cantonese.

Why Join Us:

Our client is one of the very few cloud companies that builds its own hardware from the board up. Their mission is to deliver a sovereign, energy-efficient, high-performance alternative to traditional cloud providers, and the servers you test and validate are the physical foundation of that platform.
Expect real ownership from day one, a fast-paced environment that values autonomy and problem-solving, and the chance to build something that doesn't exist yet.

Pragmatike• is dedicated to a fair, transparent, and inclusive recruitment process. We ensure that no applicant is discriminated against based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation. In accordance with the General Data Protection Regulation (GDPR), your personal data will be processed lawfully, fairly, and securely. We collect and use your personal data solely for recruitment purposes, including sharing it with our client(s) for employment consideration. You have the right to request access, correction, or deletion of your data at any time. We are committed to maintaining the confidentiality and security of your information throughout the recruitment process.

Skills Required

  • Hands-on experience testing and validating GPU/HPC servers, including bring-up, burn-in, and validation
  • Strong Linux skills at the system and hardware level
  • Firmware knowledge covering BIOS/UEFI, BMC/IPMI, GPU firmware, NIC firmware, flashing, and troubleshooting
  • Working knowledge of NVIDIA GPU drivers, CUDA, NCCL, and multi-GPU benchmarking and diagnostics
  • Understanding of server hardware, including motherboards, PCIe topology, power, cooling, storage, and networking
  • Experience with quality control or manufacturing test processes in a hardware production environment
  • Fluent English and fluent Mandarin
  • Willingness to work on-site in Hong Kong, including factory floor and laboratory environments
  • Experience with containerized or micro-datacenter deployments
  • Familiarity with InfiniBand, RoCE, or NVLink validation
  • Bash or Python scripting for test automation and reporting
  • Exposure to hardware supply chains or contract manufacturers in Hong Kong or Southern China
  • Cantonese language skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Marseille
11 Employees
Year Founded: 2022

What We Do

Trusted by remote-first companies worldwide. Completing tech projects for startups and scaleups.

Similar Jobs

Pragmatike Logo Pragmatike

Hardware Engineer

Information Technology • Software
Remote
6 Locations
11 Employees

UL Solutions Logo UL Solutions

Engineer

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
China
15000 Employees

UL Solutions Logo UL Solutions

Engineer

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
China
15000 Employees

UL Solutions Logo UL Solutions

Global EMC Operations & Strategy Manager

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
9 Locations
15000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account