Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.
The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.
RESPONSIBILITIES
Reliability Engineering
- Define and manage SLIs, SLOs, and Error Budgets for critical services
- Drive production readiness reviews and reliability requirements into architecture and design
- Perform capacity planning, failure mode analysis, and dependency risk assessments
- Identify systemic reliability risks and drive remediation before they cause customer impact
Observability
- Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
- Improve signal-to-noise ratio and reduce alert fatigue
- Build dashboards and telemetry that reflect true service health, not just infrastructure metrics
Incident Management
- Lead technical response for high-severity incidents
- Drive blameless postmortems and root cause analysis focused on systemic fixes
- Continuously improve detection, response, and recovery processes
- Participate in an on-call rotation
Automation & Toil Reduction
- Identify and eliminate manual, repetitive operational work through automation
- Build self-healing systems, tooling, and scripts to reduce human intervention
- Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
- Support Infrastructure as Code (Terraform, Bicep, or similar)
Performance & Scalability
- Conduct load testing, performance benchmarking, and bottleneck analysis
- Partner with engineering to design systems for horizontal scalability and fault tolerance
Collaboration & Culture
- Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
- Mentor engineers on SRE best practices
- Promote a culture of engineering-driven reliability over reactive operations
EDUCATION AND EXPERIENCE QUALIFICATIONS
Required Qualifications
- 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
- 2+ years experience with Kubernetes and containerized workloads
- 4-year degree in Computer Science or related field
Preferred Qualifications
- Experience with chaos engineering or resiliency testing
- Experience with high-volume, high-availability transactional systems
- Experience with AI-assisted observability or operational automation
- Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms
REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES
- Strong programming/scripting skills (Python, Go, Java, or Node.js)
- Demonstrated experience defining and operating against SLOs/Error Budgets
- Strong skills in leading incident response and root cause analysis for production systems
- Solid understanding of distributed systems and microservices architecture
- Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
- Expertise with observability platforms and monitoring strategy
This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.
Skills Required
- 5+ years of experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
- 2+ years of experience with Kubernetes and containerized workloads
- Four-year degree in Computer Science or a related field
- Strong programming or scripting skills in Python, Go, Java, or Node.js
- Experience defining and operating against SLOs and error budgets
- Experience leading incident response and root cause analysis for production systems
- Strong understanding of distributed systems and microservices architecture
- Deep expertise in at least one major cloud platform: Azure, AWS, or GCP
- Expertise with observability platforms and monitoring strategy
- Experience with chaos engineering or resiliency testing
- Experience with high-volume, high-availability transactional systems
- Experience with AI-assisted observability or operational automation
- Meaningful contributions to internal SRE tooling, frameworks, or platforms
Inspire Brands Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Inspire Brands and has not been reviewed or approved by Inspire Brands.
-
Healthcare Strength — Comprehensive health coverage, mental health support, disability insurance, and related programs are emphasized for support‑center and eligible management roles. Materials also reference options like HSAs and EAPs as part of a broad healthcare offering.
-
Leave & Time Off Breadth — Unlimited PTO is highlighted for support‑center roles alongside paid leaves such as parental and adoption assistance. Feedback suggests these time‑off elements are a core part of the corporate employee value proposition.
-
Wellbeing & Lifestyle Benefits — Employee food discounts, on‑site amenities (e.g., gym, snacks, and similar perks), and lifestyle programs add tangible day‑to‑day value. Additional benefits like commuter options, financial wellness tools, and pet insurance further broaden the package.
Inspire Brands Insights
What We Do
Inspire Brands was founded in February 2018 with a vision to invigorate great brands and supercharge their long-term growth. In an industry facing increasing disruption, our leaders saw an opportunity to build a restaurant company unlike any other – one that brings together differentiated yet complementary brands and aims to make them stronger than they would be on their own. Found inherently in the purposes of our family of brands, we identified a common thread between our restaurants – the capacity to inspire. From guest experience to career development to community well-being, Inspire plays a role in the lives of millions of people every day.








