Staff Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Richmond Hill, ON, CAN
In-Office
Senior level
Other • Security
The Role
Own escalated production incidents across application code, data pipelines, cloud infrastructure, and networks. Lead root cause analysis and permanent remediation, improve observability, manage Terraform infrastructure, operate Kubernetes across Azure and AWS, execute upgrades and migrations, strengthen change safety and automation, and support government and data-residency-sensitive environments. Participate in on-call rotations and represent Engineering during severe customer incidents.
Summary Generated by Built In

About Johnson Controls 

Johnson Controls, a global leader in thermal management, mission-critical building systems, energy efficiency, and decarbonization, helps customers use energy more productively, reduce carbon emissions, and operate with the precision and resilience required in rapidly expanding industries such as data centers, healthcare, pharmaceuticals, advanced manufacturing, and higher education.  

For more than 140 years, Johnson Controls has delivered performance where it really matters. Backed by advanced technology, lifecycle services and an industry-leading field organization, we elevate customer performance, turn goals into real-world results and help move society forward.  

Visit johnsoncontrols.com for more information and follow @Johnsoncontrols on social platforms.  

What you will do

OpenBlue from Johnson Controls is a cyber-secured smart building ecosystem that unifies data, AI, and automation to transform how buildings perform. By connecting systems that have historically stood apart and applying award-winning analytics, we give customers real-time visibility, predictive insight, and automated action across the entire building lifecycle. None of that reaches a customer without the platform underneath it. Our data platform, our enterprise SaaS portfolio, and OpenBlue Airwall run continuously for enterprise, public sector, and government customers, in environments where an outage or a data integrity problem carries real operational consequence for the buildings and the people inside them. 

Johnson Controls is seeking a Staff Site Reliability Engineer. You will be the senior technical owner of escalated production problems, capable of debugging a failure across application code, data pipelines, cloud services, and network paths, and driving it through to permanent corrective action rather than a restart and a hopeful note in the ticket. You will also plan and execute the infrastructure work that keeps those platforms healthy, with Terraform as your native language and change safety as your standing constraint. You will be a core team member of our engineering department, and participate in the on-call rotation, and occasionally join customer conversations when the severity of an issue warrants an engineer in the room. 

This role is based in Canada. We support Canadian government customers and commercial customers with Canadian data residency requirements, and this position works directly with those systems and their data. 

How you will do it

Application reliability and L3 escalation 

  • Serve as the senior escalation owner for production issues across the OpenBlue Data Platform, our enterprise SaaS products, and Airwall 

  • Debug complex, cross layer failures spanning application code, data pipelines, cloud infrastructure, and network paths 

  • Lead root cause analysis and drive both interim and permanent corrective action to closure with the owning engineering teams, including the code or configuration change that prevents recurrence 

  • Turn recurring escalations into engineering work by feeding defect patterns, reliability gaps, and supportability problems back into the product backlog 

  • Improve detection ahead of the customer by strengthening instrumentation, monitors, dashboards, and runbooks in Datadog and Grafana 

  • Participate in the on-call rotation and act as a senior technical lead during major incidents 

  • Represent Engineering directly with customers during high severity incidents and post incident reviews when the situation calls for it 

Infrastructure and platform engineering 

  • Plan and execute infrastructure upgrades, migrations, and platform changes across Azure and AWS with minimal customer disruption 

  • Own infrastructure as code in Terraform, including module design, state management, and drift remediation 

  • Operate and improve Kubernetes workloads across capacity, autoscaling, resource limits, and deployment reliability 

  • Raise release and change safety, and reduce manual toil through automation 

  • Harden the environments serving government and data residency sensitive customers, partnering with security and compliance on controls and evidence 

  • Set operational standards for observability, change management, and production readiness that other engineering teams adopt 

What you will need
Required
 

  • Must reside in Canada and be legally authorized to work in Canada without sponsorship. This requirement supports our Canadian government customers and customers with Canadian data residency obligations 

  • 7+ years in site reliability engineering, production engineering, L3 application support, or infrastructure engineering 

  • Demonstrated ability to take a complex production problem end to end and land a permanent fix, with examples you can walk through 

  • Strong hands on Terraform, with real production ownership of infrastructure as code 

  • Production experience across Azure and AWS, and with Kubernetes at scale 

  • Practical depth in Datadog and Grafana, covering instrumentation, dashboards, monitor design, and alert quality

  • Experience with AI native development, knowing when and when not to use harnesses such as Claude, Copilot, Codex, or Cursor as a core part of your workflow 

  • Willingness to participate in an on-call rotation and to respond to production incidents outside business hours 

Preferred

  • Working proficiency in Java and C#, sufficient to read, diagnose, and correct application code 

  • Experience supporting government or public sector customers, including data residency, data sovereignty, or Protected B handling requirements 

  • Eligible to obtain Government of Canada security screening at Reliability Status or higher 

  • Experience operating data platforms, including streaming and batch pipelines, data quality monitoring, and latency service level objectives 

  • Familiarity with operational technology networking and zero trust network architecture 

  • Prior experience in an L2 or L3 support organization with formal service level agreements and escalation structures 

  • Exposure to building automation, HVAC, or connected building technology 

#LI-ONSITE

Johnson Controls’ Canadian subsidiaries are committed to providing reasonable accommodation to applicants, candidates and employees with disabilities, in accordance with applicable human rights legislation, and in Ontario, in accordance with the Accessibility for Ontarians with Disabilities Act (“AODA”). When requested, accommodation will be provided throughout all stages of the recruitment and selection process. To request accommodation, please contact us. Any information you provide related to accommodation measures will be treated as confidential. A copy of Johnson Controls’ applicable AODA policies are available on our website at www.johnsoncontrols.com for your reference, and can be made available in accessible formats upon request.

Skills Required

  • Reside in Canada and be legally authorized to work in Canada without sponsorship
  • 7+ years of experience in site reliability engineering, production engineering, L3 application support, or infrastructure engineering
  • Ability to resolve complex production problems end to end and implement permanent fixes
  • Strong hands-on Terraform experience with production infrastructure-as-code ownership
  • Production experience with Azure, AWS, and Kubernetes at scale
  • Practical experience with Datadog and Grafana, including instrumentation, dashboards, monitors, and alert quality
  • Experience with AI-native development workflows and tools such as Claude, Copilot, Codex, or Cursor
  • Willingness to participate in on-call rotations and respond outside business hours
  • Working proficiency in Java and C# sufficient to read, diagnose, and correct application code
  • Experience supporting government or public sector customers, including data residency, data sovereignty, or Protected B requirements
  • Eligible to obtain Government of Canada security screening at Reliability Status or higher
  • Experience operating data platforms, including streaming and batch pipelines, data quality monitoring, and latency service level objectives
  • Familiarity with operational technology networking and zero trust network architecture
  • Prior experience in an L2 or L3 support organization with formal service-level agreements and escalation structures
  • Exposure to building automation, HVAC, or connected building technology

Johnson Controls Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Johnson Controls and has not been reviewed or approved by Johnson Controls.

  • Retirement Support Retirement support is positioned as a meaningful part of the package through employer 401(k) matching, repeatedly framed as a strong pillar of the overall rewards mix. The matching contribution is described with specific match levels in multiple places, reinforcing perceived value for long-term saving.
  • Leave & Time Off Breadth Time off is presented as comparatively robust, with multiple paid holiday categories, vacation time, and sick time described as generous or “amazing” in places. Paid time off breadth appears to be a consistent contributor to total rewards attractiveness beyond base pay.
  • Flexible Benefits Benefits are described as broad and customizable, spanning standard medical/dental/vision plus optional add-ons like pet insurance, identity protection, and legal support. Tuition reimbursement is repeatedly highlighted as a high-value option supporting professional development.

Johnson Controls Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chennai
100,000 Employees
Year Founded: 1885

What We Do

At Johnson Controls, we transform the environments where people live, work, learn and play. From optimizing building performance to improving safety and enhancing comfort, we drive the outcomes that matter most. Dedicated to protecting the environment, we deliver our promise in industries such as healthcare, education, data centers and manufacturing. With a global team of 100,000 experts in more than 150 countries and over 130 years of innovation, we are the power behind our customers’ mission. Our leading portfolio of building technology and solutions includes some of the most trusted names in the industry, such as Tyco®, York®, Metasys®, Ruskin®, Titus®, Frick®, Penn®, Sabroe®, Simplex®, Ansul® and Grinnell®.

Similar Jobs

Hybrid
2 Locations
600 Employees
Hybrid
2 Locations
600 Employees

Caseware Logo Caseware

Site Reliability Engineer

Cloud • Fintech • Software • Analytics
In-Office or Remote
Toronto, ON, CAN
570 Employees
140K-155K Annually

Enverus Logo Enverus

Site Reliability Engineer

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
2 Locations
1800 Employees

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account