Senior Site Reliability Engineer in Test, SDET

Posted Yesterday
Be an Early Applicant
Shanghai, Shanghai Municipality, Shanghai, CHN
In-Office
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Design, build, operate, and harden scalable test environments and CI/CD pipelines. Manage SBOMs and vulnerability gating, provision Kubernetes-based ephemeral and long-lived test clusters, define SLIs/SLOs and reliability practices, apply AI/ML techniques for provisioning and flaky-test detection, and collaborate with development teams to triage and automate remediation for environment and pipeline failures.
Summary Generated by Built In

At NVIDIA, we are at the forefront of technological innovation, pushing the boundaries of AI and accelerated computing. Our team in Shanghai, China is looking for a Senior Site Reliability Engineer focused on Test Environment Management to join us. This is an opportunity to build and operate highly reliable test infrastructure, CI/CD systems, and environments that power validation of NVIDIA enterprise offerings. If you are passionate about reliability engineering, test infrastructure, and AI-scale systems, this role is for you.


What you’ll be doing
  • Design, build, operate, and continuously improve reliable, scalable test environments and automation infrastructure that support validation of NVIDIA enterprise offerings.

  • Own end-to-end CI/CD pipelines using GitLab CI, GitHub Actions, and ArgoCD (GitOps) — including pipeline design, reliability, performance, and progressive delivery of test workloads.

  • Manage Software Bills of Materials (SBOMs): generation, continuous monitoring, vulnerability correlation, policy enforcement, and integration into CI/CD and release gate..

  • Provision, scale, observe, and lifecycle-manage ephemeral and long-lived test environments (Kubernetes-based and hybrid) with strong emphasis on isolation, reproducibility, and rapid recovery.

  • Define and drive reliability practices for test systems: SLIs/SLOs/error budgets for test environments and pipelines, toil reduction, chaos/resilience testing of infrastructure, and automated remediation.

  • Collaborate closely with development and platform teams to triage environment and pipeline failures, perform root-cause analysis, verify fixes, and continuously harden test infrastructure.

  • Apply AI/ML/Agentic techniques and internal tools to accelerate environment provisioning, flaky-test detection, capacity planning, anomaly detection, and overall Quality Assurance velocity.

What we need to see
  • MS or PhD in Computer Science, related field, or equivalent experience with 8+ years of professional experience in Site Reliability Engineering, Test Environment Management, CI/CD platform engineering, or software testing infrastructure.

  • Strong proficiency with Linux, shell scripting, and Python (or equivalent automation languages).

  • Hands-on experience designing and operating CI/CD systems with GitLab CI and/or GitHub Actions, practical experience with ArgoCD (or equivalent GitOps tooling) for CD of applications and infrastructure. Solid background in containerization and orchestration (Docker, Kubernetes) and virtualization technologies.

  • Deep understanding of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, toil elimination, and reliability engineering for complex distributed systems.

  • Experience building and operating test environments (ephemeral, multi-tenant, or production-like) with focus on reliability, isolation, and rapid turnaround.

  • Strong knowledge of QA principles and how test infrastructure enables high-quality software delivery.

  • Comfort working with AI/LLM-related workloads and toolings. Excellent problem-solving skills, clear written and verbal communication, and the ability to collaborate across engineering teams.

  • Self-motivated, proactive, and passionate about learning new technology at scale.

Ways to stand out from the crowd
  • Experience operating large-scale Kubernetes platforms and GitOps workflows in production or high-stakes test environments.

  • Background in software supply-chain security, SBOM tooling ecosystems, vulnerability management, and policy enforcement (OPA/Gatekeeper, Kyverno, etc.).

  • Hands-on work with NVIDIA GPU hardware, multi-GPU environments, or accelerated computing infrastructure. Experience with parallel programming, high-performance computing, or large-scale AI model training/inference test harnesses.

  • Track record of applying AI/observability techniques to detect flaky tests, optimize environment utilization, or automate root-cause analysis.

  • Prior experience defining and driving reliability programs (error budgets, chaos engineering, capacity forecasting) for CI/CD or test platforms.

NVIDIA is widely considered one of the technology world’s most desirable employers. We have some of the most brilliant people on the planet working for us. If you’re creative, autonomous, and excited about making test infrastructure as reliable as the products it validates, we want to hear from you!

Skills Required

  • MS or PhD in Computer Science or related field, or equivalent experience with 8+ years in SRE, Test Environment Management, CI/CD platform engineering, or software testing infrastructure.
  • Strong proficiency with Linux and shell scripting.
  • Proficiency in Python or equivalent automation languages.
  • Hands-on experience designing and operating CI/CD systems with GitLab CI and/or GitHub Actions.
  • Practical experience with ArgoCD or equivalent GitOps tooling for continuous delivery.
  • Solid background in containerization and orchestration (Docker, Kubernetes).
  • Experience with virtualization technologies.
  • Deep understanding of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, toil elimination.
  • Experience building and operating test environments (ephemeral, multi-tenant, production-like) focused on reliability and isolation.
  • Strong knowledge of QA principles and how test infrastructure supports software quality.
  • Comfort working with AI/LLM-related workloads and toolings.
  • Excellent problem-solving, written and verbal communication, and cross-team collaboration skills.
  • Self-motivated and proactive with passion for learning and operating technology at scale.
  • Experience operating large-scale Kubernetes platforms and GitOps workflows in production or high-stakes test environments.
  • Background in software supply-chain security, SBOM tooling ecosystems, vulnerability management, and policy enforcement (OPA/Gatekeeper, Kyverno).
  • Hands-on work with NVIDIA GPU hardware, multi-GPU or accelerated computing infrastructure; experience with parallel programming or HPC test harnesses.
  • Track record applying AI/observability techniques to detect flaky tests, optimize utilization, or automate root-cause analysis.
  • Prior experience defining and driving reliability programs (error budgets, chaos engineering, capacity forecasting) for CI/CD or test platforms.

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Airwallex Logo Airwallex

Director, Sales, SME & Growth

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office
2 Locations
2300 Employees

Airwallex Logo Airwallex

Director, Account Management, SME & Growth, CN

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office
4 Locations
2300 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Analyst, Accounts Payable

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Shanghai, Shanghai Municipality, Shanghai, CHN
16000 Employees
Hybrid
Shanghai, Shanghai Municipality, Shanghai, CHN
289097 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account