Senior Engineer, NCX

Posted Yesterday
Be an Early Applicant
4 Locations
Remote
293K-650K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Lead Day 2 operational readiness for NVIDIA Cloud Partners running large-scale GPU infrastructure. Build continuous validation, observability, automated detection and remediation, fleet lifecycle management, runbooks, SLOs, and reusable operational frameworks across compute, networking, storage, Kubernetes, and AI workloads. Collaborate directly with partner engineering and operations teams to improve infrastructure reliability, performance, and production readiness.
Summary Generated by Built In

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production.

Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations.

What you'll be doing:

  • Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation.

  • Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads.

  • Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.

  • Develop automated detection and remediation. Build workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service while minimizing disruption to customer workloads.

  • Refine fleet lifecycle administration. Implement scalable strategies for managing sizable GPU fleets, including NVIDIA driver and firmware lifecycle administration, Kubernetes node maintenance, OS patching, configuration management, upgrades, and configuration drift identification.

  • Operationalize NVIDIA reference architectures. Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable operational standards.

  • Define operational health and readiness. Develop health signals, SLOs, important metrics, acceptance criteria, and ongoing validation mechanisms that provide NVIDIA and NCPs with clear insight into infrastructure reliability and service readiness.

  • Build reusable operational frameworks. Develop tooling, automation, implementation guides, runbooks, operational playbooks, and reference implementations that can be applied consistently across multiple NCP environments.

What we need to see:

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.

  • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.

  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.

  • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.

  • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements.

  • Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.

  • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers.

  • Programming and automation experience using Python, Go, shell scripting, or similar languages.

Ways to stand out from the crowd:

  • Experience managing extensive GPU or accelerated computing infrastructure that supports AI training and inference workloads.

  • Experience with NVIDIA technologies including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.

  • Proven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency.

  • Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines and translating reference architectures or infrastructure requirements into repeatable production operating models across multiple customer or partner environments.

  • Knowledge of failure modes related to large distributed AI workloads and the infrastructure features necessary to consistently support extended training and production inference.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and dedicated people in the world working for us.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.

Skills Required

  • BS, MS, or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, a related technical field, or equivalent experience
  • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production
  • Deep understanding of Kubernetes, containers, cluster scheduling, and large multi-node environment operations
  • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and service-level agreements
  • Experience developing infrastructure lifecycle management, failure detection, remediation, upgrade, and configuration-management automation
  • Strong networking fundamentals and experience troubleshooting distributed systems across compute, network, and storage layers
  • Programming and automation experience using Python, Go, shell scripting, or similar languages
  • Experience managing large-scale GPU or accelerated-computing infrastructure for AI training and inference workloads
  • Experience with NVIDIA DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software
  • Experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or large service-provider infrastructures
  • Experience operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency
  • Knowledge of Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines
  • Knowledge of failure modes in large distributed AI workloads and infrastructure requirements for extended training and production inference

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Quantum Metric, Inc. Logo Quantum Metric, Inc.

Solutions Engineer

eCommerce • Enterprise Web • Information Technology • Software • Database • Analytics • Business Intelligence
In-Office or Remote
Madrid, Comunidad de Madrid, ESP
418 Employees
60K-80K Annually

Pfizer Logo Pfizer

Scientist

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
8 Locations
121990 Employees

Invenergy Logo Invenergy

Senior PI Administrator

Greentech • Real Estate • Social Impact • Energy • Industrial • Solar • Renewable Energy
Remote or Hybrid
18 Locations
2500 Employees
125K-155K Annually

CrowdStrike Logo CrowdStrike

Technical Sales Content Developer and Trainer for Europe (Barcelona, Spain/ Reading, UK)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
2 Locations
11000 Employees

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account