Senior Site Reliability Engineer (Onsite Role)

Posted 10 Days Ago
Be an Early Applicant
Headquarters, AZ, USA
In-Office
Senior level
Automotive • Retail
The Role
Own reliability, availability, scalability, and performance for customer-facing eCommerce microservices. Define SLOs, SLIs, and error budgets; build observability and resiliency systems; lead incident response and postmortems; automate toil and remediation; perform root-cause analysis, capacity planning, and production readiness reviews; and mentor engineers while partnering across product, security, and platform teams.
Summary Generated by Built In

Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical customer-facing microservices that power all eCommerce channels. This role applies Google-inspired SRE principles to balance feature velocity and system reliability using Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.

The role combines software engineering, cloud engineering, automation, and production operations, with a strong emphasis on building systems that are observable, resilient, and operable by default.

This is an on-site position located in Springfield, MO. Remote work is not an option for this position.

Primary Responsibilities
  • Define, implement, and own SLIs, SLOs, and error budgets for critical microservices in collaboration with product and engineering teams.
  • Use error budgets to influence release decisions, prioritize reliability initiatives, and manage operational risk.
  • Design and maintain observability platforms, including metrics, logs, traces, and real-time telemetry.
  • Track, manage, and reduce operational toil by converting repetitive operational tasks into Jira stories and epics with clear ownership and measurable outcomes.
  • Design, implement, and validate resiliency mechanisms such as graceful degradation, redundancy, automated failover, and disaster recovery.
  • Lead incident response efforts, act as an escalation point for high-severity incidents, and drive blameless postmortems.
  • Capture incident action items and reliability improvements in Jira, ensuring accountability, closure, and continuous improvement.
  • Partner with Scrum teams to improve reliability through release readiness reviews, production change validation, and testing strategies.
  • Perform deep root cause analysis, debugging, and performance tuning across distributed systems.
  • Promote shift-left reliability practices by embedding operability, monitoring, and failure testing early in the SDLC.
  • Drive continuous improvement through automation, self-healing systems, chaos engineering, and capacity planning.
  • Maintain runbooks, playbooks, and knowledge repositories, linking documentation to Jira tasks to reduce MTTR.
  • Provide technical leadership and mentoring to junior SREs and engineers.
  • Collaborate with global, distributed teams, leveraging Jira for transparent planning, dependency tracking, and execution.
  • Conduct production readiness reviews and ensure services meet operational excellence standards before deployment.
  • Track and improve operational KPIs such as availability, MTTR, MTTD, deployment success rate, and incident recurrence.
  • Collaborate with security and platform teams to ensure reliability, compliance, and operational security best practices are embedded into systems and deployment pipelines.
  • Explore opportunities to leverage AI-driven observability, anomaly detection, and operational automation to improve system reliability and reduce manual effort.
Core Competencies & Qualifications
  • 4+ years of experience in SRE, software engineering, or production operations supporting large-scale eCommerce platforms.
  • Hands-on experience with Java/J2EE-based distributed systems; React experience is a plus.
  • Proven ability to design and operate systems using SLO-driven reliability models.
  • Experience defining and measuring SLIs, including availability, latency, error rates, throughput, and saturation.
  • Good understanding of NoSQL technologies and RDBMS concepts, with the ability to write and troubleshoot database queries.
  • Experience deploying and operating services on cloud platforms such as AWS, Azure, or Google Cloud Platform (GCP).
  • Expertise with observability, APM, and caching tools such as Dynatrace, Splunk, ELK, Akamai, Quantum Metric, and Tealeaf.
  • Strong experience using Jira for backlog management, incident tracking, toil reduction initiatives, and cross-team coordination.
  • Ability to independently own services and drive reliability initiatives end-to-end.
  • Strong communication skills with the ability to influence engineering and product teams.
  • Experience participating in on-call rotations and handling critical/high-severity incidents.
Desired Skills
  • Experience building and operating microservices architectures using Spring Boot, Groovy, React, or similar technologies.
  • Strong understanding of CI/CD pipelines, release automation, and progressive delivery practices.
  • Experience working within eCommerce domains such as Catalog, Customer Data, and Order Management.
  • Familiarity with search platforms including Endeca, Solr, Lucene, and Elasticsearch.
  • Proficiency in scripting and automation using Python, Bash, Ruby, Perl, or PowerShell.
  • Experience with ITSM tools integrated with Jira workflows.
  • Exposure to capacity planning, load testing, and chaos engineering practices.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes (EKS, AKS, or GKE).
  • Familiarity with Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Ansible.
  • Understanding of operational KPIs including availability, MTTR, MTTD, deployment success rate, and incident recurrence metrics.
  • Experience conducting production readiness reviews and implementing operational governance processes.
  • Ability to collaborate with security and platform engineering teams to ensure reliability, compliance, and operational security best practices.
  • Exposure to AI-assisted operations, anomaly detection, intelligent alerting, and automated remediation solutions.
  • Experience designing scalable, self-healing platforms and automation frameworks for cloud-native environments.

O’Reilly Auto Parts has a proven track record of growth and stability. O’Reilly is full of successful career stories and believes in a strong promote-from-within philosophy, encouraging you to grow your career along with the organization. 

Total Compensation Package:

  • Competitive Wages & Paid Time Off

  • Stock Purchase Plan & 401k with Employer Contributions Starting Day One

  • Medical, Dental, & Vision Insurance with Optional Flexible Spending Account (FSA)

  • Team Member Health/Wellbeing Programs

  • Tuition Educational Assistance Programs

  • Opportunities for Career Growth

O’Reilly Auto Parts is an equal opportunity employer. The Company does not discriminate on the basis of race, religion, color, national origin or ancestry (including immigration status or citizenship), sex, sexual orientation, gender identity, pregnancy (including childbirth, lactation, and related medical conditions,) age (40 and over), veteran status, uniformed service member status, physical or mental disability, genetic information (including testing or characteristics) or another protected status as defined by local, state, or federal law, as applicable.

Qualified individuals with a disability may be entitled to reasonable accommodation under the Americans with Disabilities Act. If you require a reasonable accommodation during the application or employment process, please send an email to: [email protected] or call (800) 471-7431 option , and provide your requested accommodation, and position details.

Skills Required

  • 4+ years of experience in SRE, software engineering, or production operations supporting large-scale eCommerce platforms
  • Hands-on experience with Java/J2EE-based distributed systems
  • Experience designing and operating systems using SLO-driven reliability models
  • Experience defining and measuring SLIs such as availability, latency, error rates, throughput, and saturation
  • Understanding of NoSQL technologies and RDBMS concepts, including writing and troubleshooting database queries
  • Experience deploying and operating services on AWS, Azure, or Google Cloud Platform
  • Expertise with observability, APM, and caching tools including Dynatrace, Splunk, ELK, Akamai, Quantum Metric, or Tealeaf
  • Strong experience using Jira for backlog management, incident tracking, toil reduction, and cross-team coordination
  • Ability to independently own services and drive reliability initiatives end-to-end
  • Strong communication skills and ability to influence engineering and product teams
  • Experience participating in on-call rotations and handling critical or high-severity incidents
  • React experience
  • Experience with microservices using Spring Boot, Groovy, React, or similar technologies
  • Understanding of CI/CD pipelines, release automation, and progressive delivery
  • Experience in eCommerce domains such as Catalog, Customer Data, and Order Management
  • Familiarity with Endeca, Solr, Lucene, or Elasticsearch
  • Proficiency in Python, Bash, Ruby, Perl, or PowerShell scripting
  • Experience with ITSM tools integrated with Jira workflows
  • Exposure to capacity planning, load testing, and chaos engineering
  • Experience with Docker and Kubernetes, including EKS, AKS, or GKE
  • Familiarity with Terraform, CloudFormation, or Ansible
  • Experience conducting production readiness reviews and implementing operational governance
  • Exposure to AI-assisted operations, anomaly detection, intelligent alerting, or automated remediation
  • Experience designing scalable, self-healing cloud-native platforms and automation frameworks

O’Reilly Auto Parts Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about O’Reilly Auto Parts and has not been reviewed or approved by O’Reilly Auto Parts.

  • Equity Value & Accessibility — Equity-style upside is strengthened by an Employee Stock Purchase Plan that allows eligible full-time team members to buy ORLY shares via payroll deductions at a discount, improving access to ownership. This creates a tangible wealth-building lever that is positioned as stronger than many retail peers’ stock purchase offerings.
  • Inclusive Benefits Coverage — Retirement access is broadened because both part-time and full-time team members are immediately eligible to enroll in the 401(k). Whole-person support also extends beyond full-time staff through O’Care Solutions, which is described as available to full- and part-time team members at no cost.
  • Wellbeing & Lifestyle Benefits — Wellness and support benefits are expanded through the Live Life Well program, which can reduce medical premium share when specific health criteria are met. O’Care Solutions adds lifestyle and wellbeing coverage through counseling and practical supports like legal/financial consults and caregiving resources.

O’Reilly Auto Parts Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Springfield, MO
21,231 Employees
Year Founded: 1957

What We Do

O’Reilly Auto Parts started as a single store and has grown into a leading retailer in the automotive aftermarket industry with more than 6,100 locations and counting. With more than 94,000 team members, O’Reilly has expanded into 48 states, Puerto Rico, Mexico, and Canada. O’Reilly, headquartered in Springfield, Missouri, has a deep commitment to serving our customers, community, and our team members. Our culture values make O’Reilly the best place to work and grow! Whether you're interested in running a local store, managing a distribution center, or climbing the corporate ladder, O’Reilly has a career path in which you can truly thrive. Find out what it means to Live Green at our Fortune 500 Company and come work at the O! Mission: O'Reilly Automotive intends to be the dominant supplier of auto parts in our market areas by offering our retail customers, professional installers, and jobbers the best combination of price and quality provided with the highest possible service level.

Similar Jobs

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
10 Locations
5550 Employees
127K-249K Annually
In-Office
3 Locations
30196 Employees
112K-160K Annually

Circle Logo Circle

Investor Relations Manager

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
In-Office or Remote
24 Locations
1050 Employees
158K-191K Annually

Circle Logo Circle

Senior Software Engineer

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
In-Office or Remote
12 Locations
1050 Employees
153K-205K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Ford Energy Thumbnail
Automotive • Software • Energy • Utilities • Manufacturing • Renewable Energy
US
55 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account