Observability Engineer / Site Reliability Engineer

Posted 4 Days Ago
Be an Early Applicant
Chicago, IL, USA
In-Office
90K-100K Annually
Mid level
Cloud • Information Technology • Consulting • Design • Generative AI
We are aggressively arming companies with GenAI and Public Cloud technologies daily!
The Role
Design, scale, and maintain enterprise observability and alerting across multi-cloud and hybrid systems. Build observability stacks (Prometheus, Grafana, GCP observability), automate infrastructure with Terraform/Ansible, integrate CI/CD for containerized workloads on Kubernetes/GKE/OpenShift, optimize Linux systems, and implement SRE practices (SLIs/SLOs/Error Budgets).
Summary Generated by Built In
About Ontrac Solutions

Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.

Role Overview

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.

Key Responsibilities
  • GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required Skills & Qualifications
  • Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
  • OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
  • Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

__________________________________

Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.

Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
Get verified → https://pinpointverify.com/ontrac

Skills Required

  • Proven engineering experience within Google Cloud Platform (GCP) environments
  • Hands-on experience with Prometheus
  • Hands-on experience with Grafana
  • Experience with Google Cloud Observability suites (Cloud Logging, Cloud Monitoring, Trace, Profiler)
  • Experience with Grafana Enterprise Metrics (GEM)
  • Expert-level Linux/Unix system administration
  • Strong shell scripting skills
  • Professional coding proficiency in at least one language: Python, Go, Java, Perl, or advanced Shell
  • Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift
  • Infrastructure-as-code experience using Terraform and Ansible
  • CI/CD pipeline experience (GitHub, Harness, or similar) and integrating deployments
  • Experience implementing SRE best practices, including SLIs, SLOs, and Error Budgets
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chicago, Illinois
6 Employees
Year Founded: 2010

What We Do

Ontrac Solutions helps organizations adopt emerging technologies to scale smarter. We build GenAI platforms, predictive analytics solutions, and drive cloud adoption. We're also a HubSpot partner, supporting landing page design, website development, CRM integration, workflows, and automation. From infrastructure to marketing ops, we deliver strategy and execution that drives growth.

Similar Jobs

In-Office
5 Locations
6000 Employees
194K-267K Annually

Focused Logo Focused

Site Reliability Engineer

Artificial Intelligence • Cloud • Information Technology • Mobile • Software • Consulting
In-Office
Chicago, IL, USA
34 Employees
130K-170K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Business Development Manager

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
3 Locations
4600 Employees
158K-228K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Sales Representative

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
United States
4600 Employees
166K-240K Annually

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
LTX Thumbnail
Conversational AI • Generative AI
Jerusalem, Israel
360 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account