Senior Network Site Reliability Engineer

Posted 23 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Own and improve enterprise and data center network reliability, availability, and operational excellence. Resolve complex incidents, monitor performance, conduct root cause analyses, and implement automation to reduce toil and meet service-level objectives. Partner with architecture and deployment teams, develop observability and operational tooling, mitigate risks, collaborate on production issues, and create knowledge-base content. The role requires extensive network operations experience, automation expertise, and proficiency with networking protocols, monitoring platforms, Linux, and service management tools.
Summary Generated by Built In

We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a reliable and efficient network infrastructure.

The ideal candidate is passionate about network operations and committed to enhancing the user experience. You'll have the opportunity to solve complex network challenges using hands-on debugging and by focusing on network automation, observability, documentation, and operational excellence. This is a critical position focused on ensuring user satisfaction and brilliance in network operations.

What you'll be doing:

  • Owning the operational aspect of the network infrastructure, ensuring its high availability and reliability, actively working on network incidents and service requests.

  • Partnering with architecture and deployment teams to guarantee that new implementations are supportable and align with production standards.

  • Advocating for and implementing automation to reduce toil and improve operational efficiency. Minimizing manual operational tasks to achieve and maintain Service Level Objectives (SLOs).

  • Monitoring network performance, identifying areas for improvement, and collaborating with relevant teams to implement refinements. Proactively identifying and mitigating network risks to promote continuous improvement.

  • Collaborating with domain experts across functions to resolve production issues swiftly and effectively, ensuring customer happiness.

  • Conducting blameless postmortems and following through on Root Cause Analyses (RCAs).

  • Discovering opportunities for operational improvements and teaming up with colleagues to devise solutions that enhance excellence and sustainability in network operations.

  • Developing knowledge base articles for automation and bots.

What we need to see:

  • Educational Background: Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.

  • Work Background: A minimum of 10 years of industry practice in network operations or related fields concentrating on support, automation & site reliability engineering. Familiarity with both enterprise and the data center networks is critical.

  • Strong proficiency in network fundamentals & fixing complex network issues with expertise in network tech like TCP/UDP, IPv4/IPv6, BGP, OSPF, ISIS, VPN, L2 switching, Firewalls, Load Balancers, Data Center Network technologies, Wireless etc. Consistent track record in network operations.

  • Monitoring Tools: Familiarity with network management tools such as Prometheus, Grafana, Alert Manager, Nautobot/Netbox, BigPanda.

  • Network Automation: Experience in automating networks using frameworks such as Salt, Ansible, Python or similar.

  • Process & Service Tooling: Skills with ServiceNow, Jira & foundational knowledge of ITIL framework

  • System Administration: Knowledge of Linux system fundamentals.

  • Problem-Solving and Communication: Detailed problem-solving approach, critical thinking, coupled with good interpersonal skills and a solid grasp of ownership and drive.

Ways to stand out from the crowd:

  • Support Automation: Track record of taking operational signals through means such as SNMP, Syslog, Streaming Telemetry to solve operational challenges.

  • Platform Exposure: Experience with Mellanox/Cumulus Linux, Cisco/ Arista, Palo Alto firewalls, Versa SDWAN, Netscalers and F5 load balancers.

  • Technical Proficiency in Programming and Scripting: Competence in Python, Go, or related programming languages. Skilled in constructing intricate systems to supervise and regulate network operations, surpassing elementary script development.

  • Network Technologies: Proficient knowledge of one or more of the following technologies: VXLAN/EVPN at Scale, MPLS, RSVP, Segment Routing, SDWAN, SASE Platforms

Skills Required

  • Bachelor’s degree in Computer Science, Electrical Engineering, a related technical field, or equivalent experience
  • At least 10 years of industry experience in network operations or related support, automation, or site reliability engineering
  • Experience with enterprise and data center networks
  • Strong network fundamentals and ability to troubleshoot complex network issues
  • Expertise with TCP/UDP, IPv4/IPv6, BGP, OSPF, ISIS, VPN, Layer 2 switching, firewalls, load balancers, data center networking, and wireless
  • Familiarity with Prometheus, Grafana, Alertmanager, Nautobot or NetBox, and BigPanda
  • Network automation experience with Salt, Ansible, Python, or similar frameworks
  • Experience with ServiceNow and Jira, plus foundational ITIL knowledge
  • Knowledge of Linux system fundamentals
  • Strong problem-solving, critical-thinking, communication, ownership, and initiative skills
  • Experience using SNMP, Syslog, or streaming telemetry for support automation
  • Experience with Mellanox/Cumulus Linux, Cisco, Arista, Palo Alto firewalls, Versa SD-WAN, Netscalers, or F5 load balancers
  • Programming or scripting proficiency in Python, Go, or related languages, including complex network operations systems
  • Knowledge of VXLAN/EVPN at scale, MPLS, RSVP, segment routing, SD-WAN, or SASE platforms

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

NVIDIA Logo NVIDIA

Site Reliability Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office
3 Locations
21960 Employees

Ericsson Logo Ericsson

Security Specialist

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
88000 Employees

UL Solutions Logo UL Solutions

VP Finance, Global Business Services

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Hybrid
3 Locations
15000 Employees
250K-300K Annually

Adyen Logo Adyen

Support Engineer

Fintech • Payments • Financial Services
Easy Apply
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4771 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account