Senior DevOps Engineer, AIOps

Posted Yesterday
Be an Early Applicant
2 Locations
Hybrid
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Own the end-to-end DevOps, infrastructure, security, release, and reliability lifecycle for AI observability microservices across SaaS and on-premises environments. Build and operate Kubernetes and Helm deployments, CI/CD pipelines, infrastructure automation, databases, storage, observability, and security controls. Partner with software and AI engineers to troubleshoot distributed systems, manage incidents, improve reliability, and optimize platform performance.
Summary Generated by Built In

NVIDIA is powering the world’s most advanced AI factories, where resilient infrastructure is essential to keep accelerated computing environments running at scale. The Agentic AIOps team is building a mission-critical observability and prediction platform - delivered as both a high-scale SaaS solution and a robust on-premises deployment for NVIDIA’s largest enterprise customers. 


As a Senior DevOps Engineer, you’ll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIA’s AI infrastructure. 


What You'll Be Doing: 

  • Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations. 
  • Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints. 
  • Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory. 
  • Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity. 
  • Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations. 
  • Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows. 
  • Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening. 
  • Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes. 

What We Need to See: 

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience. 
  • 5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices. 
  • Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting. 
  • Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible. 
  • Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices. 
  • Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting. 
  • Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting. 
  • Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment. 

Ways To Stand Out From the Crowd: 

  • Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost. 
  • Familiarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage. 
  • Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana. 
  • Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers. 
  • Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain. 

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are passionate about building mission-critical systems at the frontier of AI infrastructure, we want to hear from you. 


#LI-Hybrid

Skills Required

  • Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent experience
  • 5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices
  • Hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting
  • Strong Linux administration skills
  • Proficiency in Python and Bash for automation
  • Experience with infrastructure as code and configuration tooling such as Terraform and Ansible
  • Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices
  • Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting
  • Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and alerting
  • Understanding of secure infrastructure operations and incident response
  • Demonstrated ownership, cross-functional collaboration, and prioritization
  • Experience operating AI applications, agent platforms, or LLM services
  • Familiarity with Temporal, LangGraph, Model Context Protocol, Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage
  • Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana
  • Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers
  • Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility — Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength — Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support — Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

HiBob Logo HiBob

Senior Field Deployment Engineer (FDE)

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
IL
1350 Employees

HiBob Logo HiBob

Account Executive

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
IL
1350 Employees

HiBob Logo HiBob

Controller

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
IL
1350 Employees

ServiceNow Logo ServiceNow

Security Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Petah Tikva, ISR
29000 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account