Senior Staff Site Reliability Operations Technical Lead

Posted 5 Days Ago
Be an Early Applicant
Durham, NC, USA
In-Office
184K-265K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Leads site reliability operations and Tier 3 IT support across identity, messaging, compute, endpoints, databases, networking, and datacenter infrastructure. Owns incidents, service quality, compliance, vulnerability remediation, asset management, root-cause analysis, automation, and permanent fixes. Provides technical leadership to site engineers, supports executives, coordinates with global platform teams and vendors, and represents site priorities in regional initiatives. Requires hands-on onsite work, on-call participation, and after-hours maintenance support.
Summary Generated by Built In

For over 25 years, NVIDIA has been at the forefront of transforming computer graphics, PC gaming, and accelerated computing, driven by a legacy of continuous innovation and exceptional talent! We are now bringing to bear the immense potential of AI to usher in the next era of computing, where our GPUs power the "brains" of computers, robots, and autonomous vehicles that can comprehend the world. This pioneering work demands vision, innovation, and the world's best talent. Join our diverse and supportive environment, where

We are seeking a Site Reliability Operations Technical Lead to serve as the senior technical individual contributor for reliability and support at our Durham, North Carolina site. This role owns technical service delivery locally, acts as the support point of last resort before regional and global platform teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active Directory, Exchange, database platforms, and compute infrastructure, lead the site through major incidents, and drive out the recurring problems that consume the team’s capacity. You will also lead site-level projects, represent local requirements in global initiatives, and set the technical standard the site support team works to. The successful candidate is equally comfortable running a root cause analysis, supporting an executive before an all-hands, and briefing IT leadership on site risk.

What you'll be doing:

  • Own day-to-day site operations — incidents, requests, critical issues, and support coverage — with accountability for queue health, SLA attainment, backlog, and service quality, plus site asset and inventory management across lifecycle, refresh, procurement, and compliance.

  • Serve as Tier 3 escalation owner for the site and AMER across identity (AD, hybrid Entra ID, GPO, Kerberos/LDAP, SSO, MFA), messaging (Exchange hybrid mail flow, mailbox, SMTP relay), compute (Windows, Linux, macOS, virtualization, storage, and hands-on datacenter and lab hardware), and endpoint (M365, Teams, Intune, Autopilot, imaging through migrations) driving root cause and permanent fixes rather than repeat break-fix.

  • Own endpoint compliance, vulnerability remediation, patch management, and hardening; audit readiness and evidence; and partnership with InfoSec on incident response and privileged access.

  • Drive critical issues into global platform teams and vendors with reproduction cases and diagnostic evidence through to a committed fix.

  • Act as technical lead for site SRO engineers setting standards, reviewing work, directing blocking issues, building diagnostic rigor through mentorship, and owning the site knowledge base and runbook library.

  • Serve as the primary technical contact for site IT, partnering with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding and moves, and office and lab expansions.

  • Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting; analyze ticket and reliability trends to eliminate top recurring drivers; and champion AI-driven and agentic solutions that advance SRO strategy.

  • Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums, and lead operational decisions in the manager's absence.

What we need to see:

  • 12+ years in enterprise support engineering, infrastructure, or end user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment.

  • Deep hands-on solving across Active Directory and hybrid Entra ID, Exchange hybrid, Windows and Linux server, virtualization, enterprise storage, and datacenter hardware.

  • Enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf, or equivalent), Windows 11, and the Microsoft 365 ecosystem, plus endpoint security and vulnerability remediation.

  • Database operations support and networking fundamentals — DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting.

  • Scripting and automation in Python, PowerShell, or Bash applied to real support problems, and ServiceNow or similar ITSM.

  • Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and the rigor to pursue root cause over symptom clearing.

  • Willingness to work on-site and hands-on (including lifting and moving equipment), join an on-call rotation, and support after-hours maintenance windows and cutovers.

  • Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience.

Ways to Stand Out from the crowd:

  • Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment and non-standard availability requirements.

  • Local technical lead through a site buildout, relocation, or major migration; or experience influencing global standards and tooling roadmaps for site and regional needs.

  • Executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 264,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 15, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Skills Required

  • 12+ years of experience in enterprise support engineering, infrastructure, or end user services
  • 5+ years in a senior, lead, or escalation-tier role in a multi-site environment
  • Hands-on expertise with Active Directory, hybrid Entra ID, Exchange hybrid, Windows and Linux servers, virtualization, enterprise storage, and datacenter hardware
  • Experience with enterprise endpoint management such as Intune, Autopilot, MECM/SCCM, or Jamf; Windows 11; Microsoft 365; endpoint security; and vulnerability remediation
  • Database operations support and networking fundamentals including DNS, DHCP, VLANs, wireless, firewall policy, and switch-level troubleshooting
  • Scripting and automation experience using Python, PowerShell, or Bash for support operations
  • Experience with ServiceNow or a similar IT service management platform
  • Demonstrated technical leadership without formal authority and executive-level incident communication
  • Willingness to work onsite and hands-on, including lifting and moving equipment
  • Willingness to join an on-call rotation and support after-hours maintenance windows and cutovers
  • Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent experience
  • Experience supporting engineering, laboratory, R&D, or manufacturing environments with specialized equipment
  • Experience leading site buildouts, relocations, major migrations, or influencing global standards and tooling roadmaps
  • Experience with executive support programs, AV and hybrid conference room technologies, or automation with measurable efficiency improvements

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Concord, NC, USA
16000 Employees
15-20 Hourly

Applied Systems Logo Applied Systems

Director, Product Management

Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Remote or Hybrid
United States
3116 Employees
150K-220K Annually

Applied Systems Logo Applied Systems

User Experience Designer

Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Remote or Hybrid
2 Locations
3116 Employees
120K-175K Annually

PwC Logo PwC

PwC Internal Partnership Tax Team - State & Local Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
68 Locations
370000 Employees
212K-244K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account