Senior Staff Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Lead technical strategy and roadmaps for large-scale SRE initiatives. Design and build resilient distributed and GPU-accelerated systems. Improve automation, observability, and AI/LLM-aware monitoring. Build autonomous incident-response pipelines, collaborate with Cloud, Platform, Security, and AI/ML teams, analyze Kubernetes-scale infrastructure, and mentor engineers to adopt AI-assisted workflows.
Summary Generated by Built In

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

What you'll be doing:

  • Lead the technical strategy and roadmap for large-scale, multi-functional SRE initiatives that boost reliability, scalability, and developer efficiency throughout enterprise systems.

  • Design and build resilient distributed systems that power NVIDIA's next-generation AI-powered enterprise products and services, including transforming legacy applications and database systems into modern and scalable architectures.

  • Drive automation and observability improvements, using metrics and analytics — including AI workload quality signals and model performance telemetry — to improve performance, reliability, and efficiency.

  • Build LLM-aware monitoring and autonomous incident response pipelines to reduce toil, accelerate MTTR, and evolve on-call operations toward AI-assisted remediation.

  • Work together with Cloud, Platform, Security, and AI/ML groups to develop modern SRE elements and AI-native platform features that guarantee high availability and secure operations.

  • Analyze and run complex systems — including Kubernetes-scale and AI/ML infrastructure challenges — championing standards in system design and incident management.

  • Drive AI-assisted and AI-first engineering practices across the organization, mentoring engineers to adopt agentic development workflows, coding agents, and LLM-powered tooling in their day-to-day work.

What we need to see:

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.

  • BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience.

  • Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code and distributed systems debugging.

  • Experience with infrastructure-as-code tooling such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.

  • Solid understanding of OpenTelemetry or other observability implementations at scale, including observability build for AI workloads (model performance, drift detection, and AI quality signals).

  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP), including bare-metal and GPU-accelerated infrastructure.

  • Outstanding problem-solving, communication, and collaboration skills, with the ability to influence across technical and interpersonal boundaries.

Ways to stand out from the crowd:

  • Experience with Public Cloud or large-scale automation systems.

  • Capability to steer technical strategy and achieve quantifiable reliability results in complex, multi-team settings.

  • Experience building or operating agentic AI platforms — autonomous or semi-autonomous — including LLM toolchains, agent orchestration frameworks (e.g., LangGraph, AutoGen), or automated runbook and toil-reduction processes.

  • A strong sense of ownership, curiosity, and innovation — and turn challenges into opportunities.

NVIDIA leads the charge in innovative breakthroughs in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, functions as the visual cortex of today’s computers and forms the core of our products and services. Our work opens new realms to explore, encourages outstanding creativity and discovery, and powers inventions once thought of as science fiction — from artificial intelligence to autonomous systems. NVIDIA is searching for outstanding talent like you to help us advance the next wave of artificial intelligence!

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com

Skills Required

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles
  • BS degree in Computer Science or related technical field, or equivalent experience
  • Proficiency in Python, TypeScript, JavaScript, or Go with focus on automation and distributed systems debugging
  • Experience with infrastructure-as-code tooling (AWS CDK, CloudFormation, Terraform, or CrossPlane)
  • Understanding of OpenTelemetry or other observability implementations at scale, including AI workload observability (model performance, drift detection)
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP), including bare-metal and GPU-accelerated infrastructure
  • Outstanding problem-solving, communication, and collaboration skills
  • Experience with public cloud or large-scale automation systems
  • Experience building or operating agentic AI platforms, LLM toolchains, or agent orchestration frameworks (e.g., LangGraph, AutoGen)
  • Proven ability to steer technical strategy and achieve measurable reliability improvements across multi-team environments
  • Strong sense of ownership, curiosity, and innovation

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

OneTrust Logo OneTrust

Staff Software Engineer

Artificial Intelligence • Cloud • Information Technology • Security • Social Impact • Software • Cybersecurity
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
2000 Employees
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
528 Employees

Optum Logo Optum

Principal Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Optum Logo Optum

Machine Learning Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account