Staff System Engineer – AI Infrastructure

Posted Yesterday
Be an Early Applicant
Singapore, SGP
In-Office
Senior level
Artificial Intelligence • Cloud • Software • Big Data Analytics
Shape the Global Future of Enterprise AI Come build the future of enterprise AI with a global team that puts its people
The Role
Designs, investigates, and optimizes infrastructure for production AI workloads across GPU environments, on-premises datacenters, private infrastructure, and public clouds. Diagnoses Linux, GPU, container, networking, storage, and hardware issues; develops benchmarking, diagnostic, deployment, and validation tooling; improves inference performance, GPU utilization, reliability, and scalability; and translates technical findings into product capabilities in collaboration with engineering and product teams.
Summary Generated by Built In

Business Area:

Professional Services

Seniority Level:

Mid-Senior level

Job Description: 

At Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we're the preferred data partner for the top companies in almost every industry.  Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.

As adoption of AI grows across our data and AI platform, customers are using Cloudera AI for increasingly diverse workloads across on-premises, private, sovereign, and public-cloud environments. We are seeking a Staff Systems Engineer – AI Infrastructure to understand the infrastructure requirements and constraints behind these workloads and drive hands-on engineering solutions that improve our product.

This is a hands-on system engineering role. You will investigate ambiguous technical scenarios, identify underlying systems constraints, and develop and validate solutions through experimentation, prototyping, benchmarking, and engineering work. Success means turning diverse workload and infrastructure scenarios into validated technical solutions and, where appropriate, scalable product capabilities and improvements.

As a Staff System Engineer, you will:

  • Work across AI Infrastructure & Systems Engineering: Investigate and design solutions for AI workloads across heterogeneous GPU environments, on-premises datacenters, private infrastructure, and public clouds; reproduce complex scenarios and validate solutions through hands-on experimentation and proof-of-concepts.

  • Work across AI Workload & Inference Engineering: Develop and optimize infrastructure solutions for production AI workloads, including inference and serving, considering workload characteristics, performance, GPU capacity, resource utilization, and deployment constraints.

  • Systems & Performance Engineering: Diagnose issues across Linux, GPU runtimes and drivers, containers, networking, storage, and hardware; identify root causes and validate solutions to performance, reliability, and scalability challenges.

  • GPU Resource Efficiency: Investigate approaches for efficiently allocating and utilizing GPU resources across AI workloads, including workload-aware sharing and partitioning where appropriate.

  • Infrastructure Tooling & Validation: Build diagnostic, benchmarking, deployment, and validation tooling to reproduce complex infrastructure scenarios and evaluate product performance across different environments.

  • Product & Engineering Collaboration: Translate infrastructure findings into technical requirements, product improvements, performance optimizations, and reusable platform capabilities in partnership with product and engineering teams.

We’re excited about you if you have:

  • 8+ years of experience in systems software, distributed infrastructure, platform engineering, performance engineering, or a related field, with a track record of independently solving complex, ambiguous engineering problems.

  • Hands-on experience with production GPU-based infrastructure supporting AI workloads, with a strong understanding of the infrastructure characteristics and constraints that affect them.

  • Ability to take ambiguous problems, develop hypotheses, investigate root causes, and build or validate solutions through experimentation, debugging, prototyping, and measurement without requiring step-by-step direction.

  • Strong understanding of distributed systems, Linux, containers, and production infrastructure, with the ability to reason across multiple layers of the technology stack.

  • Demonstrated ability to diagnose and optimize bottlenecks involving GPU utilization, compute, memory, networking, I/O, or workload/runtime behavior.

  • Hands-on experience designing or operating production infrastructure in on-premises, private-cloud, and/or public-cloud environments, with an understanding of the practical constraints of heterogeneous environments.

You may also have:

  • Experience with modern AI inference and serving technologies such as NVIDIA NIM, vLLM, SGLang, Triton, or equivalent.

  • Experience with Kubernetes or distributed AI workload orchestration.

  • Experience with GPU resource management, sharing, or partitioning, including technologies such as NVIDIA MIG.

  • Experience diagnosing high-performance GPU networking or distributed communication issues.

  • Experience building infrastructure diagnostics, benchmarks, or proof-of-concept systems to evaluate new architectures or technologies.

What you can expect from us:

  • Generous PTO Policy 

  • Support work life balance with Unplugged Days

  • Flexible WFH Policy 

  • Mental & Physical Wellness programs 

  • Phone and Internet Reimbursement program 

  • Access to Continued Career Development 

  • Comprehensive Benefits and Competitive Packages 

  • Paid Volunteer Time

  • Employee Resource Groups

EEO/VEVRAA

#LI-RC1

Skills Required

  • 8+ years of experience in systems software, distributed infrastructure, platform engineering, performance engineering, or a related field
  • Hands-on experience with production GPU-based infrastructure supporting AI workloads
  • Ability to independently investigate ambiguous problems, debug root causes, prototype solutions, and validate results through experimentation and measurement
  • Strong understanding of distributed systems, Linux, containers, and production infrastructure
  • Experience diagnosing and optimizing GPU utilization, compute, memory, networking, I/O, or workload/runtime bottlenecks
  • Hands-on experience designing or operating production infrastructure in on-premises, private-cloud, and/or public-cloud environments
  • Experience with AI inference and serving technologies such as NVIDIA NIM, vLLM, SGLang, Triton, or equivalent
  • Experience with Kubernetes or distributed AI workload orchestration
  • Experience with GPU resource management, sharing, or partitioning, including NVIDIA MIG
  • Experience diagnosing high-performance GPU networking or distributed communication issues
  • Experience building infrastructure diagnostics, benchmarks, or proof-of-concept systems

Cloudera Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cloudera and has not been reviewed or approved by Cloudera.

  • Fair & Transparent Compensation Pay practices are presented as equity-audited with a recognized Fair Pay Workplace certification and ongoing internal reviews. Compensation is often characterized as competitive for similar-sized peers, with visible market-aligned ranges for key roles.
  • Healthcare Strength Benefit descriptions emphasize comprehensive medical, dental, and vision coverage, alongside life and disability insurance, an EAP, wellness programming, and U.S. gym reimbursement. Health coverage is described as strong in practice.
  • Leave & Time Off Breadth Policies include generous PTO and holidays plus recurring companywide Unplugged Days that create extended weekends. Parental and medical leave are also highlighted as part of the core package.

Cloudera Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
3,092 Employees
Year Founded: 2008

What We Do

At Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we're the preferred data partner for the top companies in almost every industry. Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.

Why Work With Us

Impact at Scale: The infrastructure we build solves massive problems for top global banks, telecommunications giants, and healthcare providers. Cutting-Edge Tech: Work directly at the intersection of Open Source, Machine Learning, and Generative AI. Unmatched Flexibility: Enjoy a remote-friendly, hybrid culture that respects your time—including

Similar Jobs

Cloudflare Logo Cloudflare

Senior Revenue Operations Support Manager, Singapore

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Cloudflare Logo Cloudflare

Senior Solutions Architect

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Cloudflare Logo Cloudflare

Solutions Architect

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Mastercard Logo Mastercard

Software Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Singapore, SGP
38800 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account