Senior Site Reliability Engineer

Posted 3 Days Ago
Be an Early Applicant
Centro, Maripí, Boyacá, COL
In-Office
Senior level
Food
The Role
Build and scale reliable, resilient, observable systems for high-traffic digital platforms. Responsibilities include defining SLOs and error budgets, improving observability, leading incident response and postmortems, automating operational work, supporting infrastructure as code and CI/CD, conducting performance analysis, and implementing scalability and resiliency patterns. The role requires mentoring engineers and participating in on-call rotations, with 80% on-site presence in Atlanta.
Summary Generated by Built In

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.

The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.

RESPONSIBILITIES

Reliability Engineering

  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact

Observability

  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics

Incident Management

  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on-call rotation

Automation & Toil Reduction

  • Identify and eliminate manual, repetitive operational work through automation
  • Build self-healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)

Performance & Scalability

  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance

Collaboration & Culture

  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering-driven reliability over reactive operations

EDUCATION AND EXPERIENCE QUALIFICATIONS

Required Qualifications 

  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Science or related field

Preferred Qualifications 

  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms

REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES

  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.


 

Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.

Skills Required

  • 5+ years of experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years of experience with Kubernetes and containerized workloads
  • Four-year degree in Computer Science or a related field
  • Strong programming or scripting skills in Python, Go, Java, or Node.js
  • Experience defining and operating against SLOs and error budgets
  • Experience leading incident response and root cause analysis for production systems
  • Strong understanding of distributed systems and microservices architecture
  • Deep expertise in at least one major cloud platform: Azure, AWS, or GCP
  • Expertise with observability platforms and monitoring strategy
  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Meaningful contributions to internal SRE tooling, frameworks, or platforms

Inspire Brands Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Inspire Brands and has not been reviewed or approved by Inspire Brands.

  • Healthcare Strength Comprehensive health coverage, mental health support, disability insurance, and related programs are emphasized for support‑center and eligible management roles. Materials also reference options like HSAs and EAPs as part of a broad healthcare offering.
  • Leave & Time Off Breadth Unlimited PTO is highlighted for support‑center roles alongside paid leaves such as parental and adoption assistance. Feedback suggests these time‑off elements are a core part of the corporate employee value proposition.
  • Wellbeing & Lifestyle Benefits Employee food discounts, on‑site amenities (e.g., gym, snacks, and similar perks), and lifestyle programs add tangible day‑to‑day value. Additional benefits like commuter options, financial wellness tools, and pet insurance further broaden the package.

Inspire Brands Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Atlanta, GA
Year Founded: 2018

What We Do

Inspire Brands was founded in February 2018 with a vision to invigorate great brands and supercharge their long-term growth. In an industry facing increasing disruption, our leaders saw an opportunity to build a restaurant company unlike any other – one that brings together differentiated yet complementary brands and aims to make them stronger than they would be on their own. Found inherently in the purposes of our family of brands, we identified a common thread between our restaurants – the capacity to inspire. From guest experience to career development to community well-being, Inspire plays a role in the lives of millions of people every day.

Similar Jobs

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

Cloudflare Logo Cloudflare

Senior Customer Engineer, LATAM - MCR Bogotá, Colombia.

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Colombia
4400 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sr. Sales Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

Domino Data Lab Logo Domino Data Lab

Support Engineer

Artificial Intelligence • Machine Learning
Remote or Hybrid
10 Locations
200 Employees

Similar Companies Hiring

Munchkin, Inc. Thumbnail
Consumer Web • eCommerce • Food • Kids + Family • Design • Manufacturing
Milton, Ontario
325 Employees
Tastewise Thumbnail
Artificial Intelligence • Big Data • Food • Retail • Software • Generative AI • Big Data Analytics
NYC, NYC
120 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account