Site Reliability Engineer-Devops with Networking(DNS, TCP/IP)

Posted 5 Days Ago
Be an Early Applicant
Sadar Bazar New Area, Ambala, Haryana, IND
In-Office
Mid level
Professional Services • Consulting
The Role
Maintain highly available and reliable production systems through automation, monitoring, alerting, logging, incident response, and performance optimization. Define and track SLIs and SLOs, perform root cause analysis, manage change and disaster recovery processes, and maintain operational documentation. Collaborate with engineering, product, operations, vendors, and cross-functional stakeholders while using cloud platforms, Docker, Kubernetes, CI/CD tools, ServiceNow, JIRA, and observability solutions.
Summary Generated by Built In
  •  Title : Site Reliability Engineer
  • Experience : 3+ years
  • Location : Gurgaon
  • Employment Type : Full Time
  • Notice Period : Immediate Joiner
  •  

 Job Description

About the Role

We are seeking a proactive and detail-oriented Site Reliability Engineer (SRE) with 3+ years of experience to ensure high availability, reliability, and performance of production systems.

This role focuses on automation, observability, incident management, and cross-team coordination to drive operational excellence.

 

Key Responsibilities

                    Maintain reliable, scalable, and secure production environments.

                    Implement and manage monitoring, alerting, and logging solutions.

                    Contribute to defining and tracking SLIs/SLOs and support error budget practices.

                    Automate operational tasks to improve efficiency and reduce manual effort.

                    Perform troubleshooting and Root Cause Analysis (RCA) for production incidents.

                    Optimize system performance, availability, and capacity.

                    Maintain run books, SOPs, and incident documentation in Confluence.

                    Adhere to change management, deployment governance, and disaster recovery standards.

                    Support incident response for critical production services.

 

Collaboration & Tools

                    Coordinate with external vendors and internal cross-functional teams.

                    Work closely with Engineering, Product Owners, and Operations teams.

                    Manage incidents and changes using ServiceNow & JIRA.

                    Collaborate through Slack and structured communication channels.



Technical Skills

Systems & Clouds

                    Strong knowledge of Windows and Linux/Unix systems.

                    Solid understanding of networking fundamentals (DNS, TCP/IP, Load Balancing, Firewalls).

                    Experience with at least one cloud platform (AWS, Azure, or GCP).

                    Automation & CI/CD

                    Proficiency in one scripting/programming language (Python, Go, Bash, PowerShell, or Java).

                    Understanding of CI/CD pipelines and automation practices.

 

Containers & Observability

                    Hands-on experience with Docker and Kubernetes.

                    Experience with monitoring tools such as Grafana or Power BI.

                    Ability to analyze logs, metrics, and traces for troubleshooting.

 

ITSM & Documentation

                    Experience with ServiceNow & JIRA (incident/change/problem workflows).

                    Working knowledge of Confluence for technical documentation and knowledge management.

Additional Experience (Preferred)

                    Background in DevOps, Cloud Engineering, or Platform Engineering

                    Understanding of security best practices and compliance standards.

                    Familiarity with AI-assisted engineering tools (Claude Code, Jellyfish, GitHub Copilot).

                    Exposure to large-scale or production-grade systems.

 

Soft Skills

                    Strong analytical and troubleshooting mindset

                    Excellent written and verbal communication skills

                    Effective stakeholder and vendor coordination

                    Ownership driven and composed during high level severity incidents


Skills Required

  • 3+ years of experience
  • Immediate availability to join
  • Strong knowledge of Windows and Linux/Unix systems
  • Understanding of DNS, TCP/IP, load balancing, and firewalls
  • Experience with at least one cloud platform: AWS, Azure, or GCP
  • Proficiency in one scripting or programming language: Python, Go, Bash, PowerShell, or Java
  • Understanding of CI/CD pipelines and automation practices
  • Hands-on experience with Docker and Kubernetes
  • Experience with monitoring tools such as Grafana or Power BI
  • Experience with ServiceNow and JIRA incident, change, and problem workflows
  • Working knowledge of Confluence for technical documentation
  • Background in DevOps, Cloud Engineering, or Platform Engineering
  • Understanding of security best practices and compliance standards
  • Familiarity with Claude Code, Jellyfish, or GitHub Copilot
  • Exposure to large-scale or production-grade systems
  • Strong analytical and troubleshooting skills
  • Excellent written and verbal communication skills
  • Stakeholder and vendor coordination skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
255 Employees
Year Founded: 2009

What We Do

Right Advisors Private Limited is a Faridabad-based human resource consulting and staffing organization serving businesses across multiple industries. Its services include recruitment, contract staffing, payroll management, executive search, recruitment process outsourcing, and workforce management. The company uses a solutions-based consulting approach and flexible staffing models to help clients improve productivity, build high-performance teams, and focus on their core operations.

Similar Jobs

The Aerospace Corporation Logo The Aerospace Corporation

Sales Manager

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
India
4600 Employees

The Aerospace Corporation Logo The Aerospace Corporation

Advanced Software Engr

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
India
4600 Employees

ServiceNow Logo ServiceNow

Sr Business Strategy Specialist

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Gurugram, Haryana, IND
29000 Employees

Ericsson Logo Ericsson

Machine Learning Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Hybrid
4 Locations
88000 Employees

Similar Companies Hiring

Quantum Rise Thumbnail
Software • Professional Services • Natural Language Processing • Machine Learning • Consulting • Automation • Artificial Intelligence
Chicago, Illinois
20 Employees
MetTel Thumbnail
Information Technology • Consulting
New York, New York
695 Employees
Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account