Senior DevOps Engineer

Posted Yesterday
Be an Early Applicant
Santa Clara, CA, USA
In-Office
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Lead the design, automation, and operation of scalable Linux infrastructure supporting networking software development and testing. Build infrastructure-as-code, configuration management, self-service tools, monitoring, and reliability practices across physical servers, networks, virtualization, containers, and storage. Diagnose complex hardware-to-application issues, establish operational standards, collaborate across teams, and mentor engineers while driving infrastructure initiatives to completion.
Summary Generated by Built In

As a Senior DevOps Engineer, you will help lead the evolution of infrastructure operations within our Networking Software group. Building on a strong Linux systems administration foundation, you will build, automate, and operate scalable platforms that support networking software development and testing. This role offers an outstanding opportunity to work with elite technology and collaborate with ambitious engineers across global sites. If you are passionate about automation, reliability, technical leadership, and continuous improvement, this is the perfect opportunity for you!

What you'll be doing:

  • Build, provision, configure, and maintain scalable Linux infrastructure for networking feature creation and validation, including physical servers, network switches, virtualization platforms, containers, and remote-management interfaces.

  • Develop automation for infrastructure provisioning, configuration management, software deployment, upgrades, and day-to-day operations using infrastructure-as-code and configuration-management practices.

  • Build reusable tools and self-service capabilities that simplify infrastructure operations, improve engineering efficiency, and reduce repetitive manual work.

  • Diagnose and resolve complex issues spanning hardware, firmware, operating systems, virtualization, containers, storage, network communications, and application environments.

  • Implement monitoring, observability, capacity management, and reliability practices to improve infrastructure performance, availability, and operational readiness.

  • Partner with engineering, IT, facilities, security, and network teams to define technical standards, maintain documentation and runbooks, and establish scalable operational processes.

  • Provide technical leadership, guide infrastructure initiatives, and mentor team members in automation, troubleshooting, and operational guidelines.

What we need to see:

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.

  • 6+ years of experience in systems engineering, DevOps, site reliability engineering, or infrastructure operations, including significant hands-on experience coordinating production or engineering Linux environments.

  • Experience working in the semiconductor industry or a hardware-focused engineering environment, with deep hands-on expertise in bare-metal Linux systems and server components—including CPUs, GPUs, memory, PCIe devices, NICs, storage, BIOS/UEFI, BMC/IPMI/Redfish, power, and cooling.

  • Proficiency in generative AI tools and skill in applying them effectively to automation, troubleshooting, documentation, operational analysis, and engineering efficiency.

  • Skilled at diagnosing complex issues across hardware, firmware, and operating-system layers.

  • Strong data-center networking knowledge, including TCP/IP, DNS, DHCP, VLANs, routing, switching, firewalls, and network troubleshooting tools.

  • Strong analytical, problem-solving, written communication, and cross-departmental collaboration skills, with the ability to guide technical initiatives to completion.

Ways to stand out from the crowd:

  • Experience managing Linux KVM/QEMU virtualization, Kubernetes clusters, multi-user engineering lab environments, NFS or distributed storage systems, and automated OS or cluster provisioning platforms.

  • Experience supporting fast-growing engineering labs, large-scale data-center environments, or globally distributed infrastructure.

  • Familiarity with observability platforms, including metrics, logging, tracing, alerting, incident management, and service-level objectives.

  • Experience crafting self-service infrastructure platforms and reusable automation that improves developer efficiency and theaAbility to establish clear, reliable, and scalable engineering and operational practices from evolving requirements.

  • Proven technical leadership, team leadership, or managerial experience, including mentoring engineers, prioritizing work, coordinating cross-functional initiatives, and driving projects from planning through completion.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 21, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience
  • 6+ years of experience in systems engineering, DevOps, site reliability engineering, or infrastructure operations
  • Hands-on experience coordinating production or engineering Linux environments
  • Experience in the semiconductor industry or a hardware-focused engineering environment
  • Deep hands-on expertise with bare-metal Linux systems and server components, including CPUs, GPUs, memory, PCIe devices, NICs, storage, BIOS/UEFI, BMC/IPMI/Redfish, power, and cooling
  • Proficiency in generative AI tools and applying them to automation, troubleshooting, documentation, operational analysis, and engineering efficiency
  • Experience diagnosing complex issues across hardware, firmware, and operating-system layers
  • Strong data-center networking knowledge, including TCP/IP, DNS, DHCP, VLANs, routing, switching, firewalls, and network troubleshooting tools
  • Strong analytical, problem-solving, written communication, and cross-departmental collaboration skills
  • Ability to guide technical initiatives to completion
  • Experience managing Linux KVM/QEMU virtualization
  • Experience managing Kubernetes clusters
  • Experience with multi-user engineering lab environments
  • Experience with NFS or distributed storage systems
  • Experience with automated OS or cluster provisioning platforms
  • Experience supporting fast-growing engineering labs, large-scale data-center environments, or globally distributed infrastructure
  • Familiarity with observability platforms, including metrics, logging, tracing, alerting, incident management, and service-level objectives
  • Experience building self-service infrastructure platforms and reusable automation
  • Technical leadership, team leadership, or managerial experience, including mentoring engineers and coordinating cross-functional initiatives

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

True Anomaly Logo True Anomaly

Senior Devops Engineer

Aerospace • Artificial Intelligence • Hardware • Machine Learning • Software • Defense • Manufacturing
In-Office
2 Locations
300 Employees
150K-225K Annually

FloQast Logo FloQast

Senior Devops Engineer

Artificial Intelligence • Fintech • Software
Hybrid
San Jose, CA, USA
800 Employees
186K-282K Annually

PayZen Logo PayZen

Senior Devops Engineer

Fintech • Healthtech • Payments
Hybrid
San Francisco, CA, USA
111 Employees
179K-210K Annually

Faro Health Inc. Logo Faro Health Inc.

Senior Devops Engineer

Cloud • Healthtech • Information Technology
In-Office or Remote
2 Locations
36 Employees
174K-205K Annually

Similar Companies Hiring

Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account