Site Reliability Engr II

Posted 24 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Aerospace • Security • Energy • Industrial
The Role
Design, build, and maintain highly available distributed systems and SRE practices. Implement monitoring, incident response, automation, IaC, and cloud infrastructure management. Lead capacity planning, DR, observability improvements, and mentor junior engineers while partnering with development teams to improve reliability and reduce operational toil.
Summary Generated by Built In

Job Title: Senior Site Reliability Engineer 

Location: Bangalore

We are seeking a highly technical SRE Engineer to design, build, and maintain fault-tolerant, scalable, and highly available distributed systems. You will champion SRE best practices, reduce manual operations (toil) via automation, and partner with product development squads to embed reliability into the software delivery lifecycle

Your role will include overseeing, supervising and reviewing tasks performed by team members to ensure effective execution of work; managing end-to-end processes and projects for both internal and external clients with responsibility for timely and accurate delivery; issuing clear instructions and directions to team  members on tasks to be performed; and mentoring and guiding junior colleagues to Support their skill development, professional growth, and overall success.

Responsibilities

Key Responsibilities:

Reliability & Availability

  • Ensure high availability and uptime of production services.
  • Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs).
  • Design and implement disaster recovery (DR) strategies, including RTO/RPO targets.
  • Conduct capacity planning and scalability assessments.

Monitoring & Observability

  • Implement and maintain monitoring, logging, and alerting systems.
  • Create dashboards to track system health and performance.
  • Improve observability using tools such as Prometheus, Grafana, Azure Monitor, Dynatrace, or Elastic.
  • Proactively detect, investigate, and resolve system issues.
  • Incident Management
  • Participate in on-call rotations and incident response.
  • Lead troubleshooting during service disruptions.
  • Perform Root Cause Analysis (RCA) and drive corrective actions.
  • Reduce Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR).

Automation & Engineering

  • Develop automation to eliminate repetitive operational tasks.
  • Build self-healing and auto-scaling capabilities.
  • Create and maintain Infrastructure as Code (IaC).
  • Improve CI/CD pipelines and deployment reliability. 

Cloud & Infrastructure Management

  • Manage cloud environments and production infrastructure.
  • Support Kubernetes/AKS clusters, databases, networking, storage, and messaging services.
  • Optimize infrastructure costs and resource utilization.
  • Ensure security, compliance, and operational best practices. 

Collaboration

  • Partner with software development teams throughout the software lifecycle.
  • Participate in architecture reviews and non-functional requirement (NFR) assessments.
  • Review reliability risks and recommend improvements.
  • Support release planning and production readiness reviews. 

Required Technical Skills

Infrastructure & Cloud

  • Azure/AWS or GCP
  • Kubernetes / AKS / OpenShift
  • Linux administration
  • Networking fundamentals (TCP/IP, DNS, Load Balancing)
  • Database administration basics (SQL/NoSQL)
  • Automation & DevOps
    • Terraform, ARM, Bicep, or CloudFormation
  • CI/CD tools (Azure DevOps, GitHub Actions, Jenkins, GitLab)
  • Infrastructure as Code (IaC)
  • Configuration management tools such as Ansible
  • Programming such as Python, Jave, Sheel scripting, Go

Monitoring & Observability

  • Grafana
  • Prometheus
  • Dynatrace
  • Elastic Stack
  • Azure Monitor / Log Analytics 
Qualifications

Experience Level: 5+ yrs

Required Qualifications

  • Education: Bachelor's/Master’s degree in Computer Science, Information Technology, or equivalent practical experience.
  • Experience supporting large-scale cloud-native applications.
  • Experience with incident management and production support.
  • Understanding of SRE concepts such as:
    • Error Budgets
    • SLI/SLO/SLA
    • Chaos Engineering
    • High Availability
    • Reliability Engineering 

Success Metrics

  • Service Availability (% Uptime)
  • SLA/SLO Compliance
  • MTTR / MTTD
  • Deployment Success Rate
  • Production Incident Reduction
  • Operational Cost Optimization
  • Automation Coverage
  • Customer Experience Metrics

About UsHoneywell Technologies is a global, pure-play automation company with a legacy of innovating to help solve the world’s most mission-critical challenges, enhancing the quality of life for people and communities around the world. We serve the building, industrial and process sectors with a broad portfolio of services, solutions and products, underpinned by our Honeywell Technologies Accelerator operating system and Honeywell Technologies Forge intelligence layer. By combining the deep domain expertise of our more than 50,000 employees with decades of data from our global installed base, we are uniquely positioned to lead the industrial sector’s transition from automation to autonomy.

Skills Required

  • 5+ years experience supporting large-scale cloud-native applications and SRE functions
  • Bachelor's or Master's degree in Computer Science, IT, or equivalent practical experience
  • Experience with incident management, on-call rotations, and production support
  • Understanding of SRE concepts (Error Budgets, SLI/SLO/SLA, Chaos Engineering, HA)
  • Azure, AWS, or GCP cloud experience
  • Kubernetes / AKS / OpenShift experience
  • Linux administration experience
  • Networking fundamentals (TCP/IP, DNS, Load Balancing)
  • Database administration basics (SQL and NoSQL)
  • Infrastructure as Code: Terraform, ARM, Bicep, or CloudFormation
  • CI/CD tooling experience (Azure DevOps, GitHub Actions, Jenkins, GitLab)
  • Configuration management such as Ansible
  • Programming/scripting: Python, Java, Shell scripting, Go
  • Monitoring and observability tools: Prometheus, Grafana, Dynatrace, Elastic Stack, Azure Monitor/Log Analytics

Honeywell Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Honeywell and has not been reviewed or approved by Honeywell.

  • Retirement Support Retirement benefits are anchored by a strong 401(k) match with clear vesting and annual funding mechanics. Plan administration and education resources further reinforce long‑term savings support.
  • Leave & Time Off Breadth Time away provisions include company holidays, flexible vacation for many exempt roles, and paid sick time. These policies contribute meaningful breadth beyond base pay.
  • Parental & Family Support Paid parental leave is available to all parents with flexible usage options, and certain family‑building supports are included. Birth mothers can coordinate leave with short‑term disability for extended coverage.

Honeywell Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Charlotte, NC
110,269 Employees
Year Founded: 1906

What We Do

Honeywell is a Fortune 500 company that invents and manufactures technologies to address tough challenges linked to global macrotrends such as safety, security, and energy. With approximately 110,000 employees worldwide, including more than 19,000 engineers and scientists, we have an unrelenting focus on quality, delivery, value, and technology in everything we make and do.

Similar Jobs

In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
10000 Employees

TransUnion Logo TransUnion

Java Engineer

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
13000 Employees

Ericsson Logo Ericsson

5G Baseband Developer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
88000 Employees

Ericsson Logo Ericsson

Solution Integrator

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
6 Locations
88000 Employees

Similar Companies Hiring

Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account