Site Reliability Engineer (SRE)

Posted 4 Days Ago
Be an Early Applicant
12 Locations
In-Office or Remote
Mid level
Information Technology
The Role
Monitor and maintain platform reliability across multiple e-commerce client environments. Respond to incidents, perform technical triage, maintain observability dashboards and alerts, track SLIs and SLOs, document runbooks, support root cause analysis, and participate in postmortems and on-call rotations. Identify recurring issues and automate operational processes while collaborating with Service Desk and EMEA teams during shift handoffs.
Summary Generated by Built In
ABOUT APPLY
 
APPLY is the Agentic Customer Experience (ACx) partner for the world's most ambitious consumer and entertainment brands. We bring together deep domain expertise across Retail, CPG, Sports, and Media with AI-native delivery capability, designing and delivering agentic solutions that turn CX vision into commercial reality. We are the partner of choice for brands like Arc'teryx, NFL, Lululemon, and Kraft Heinz. For more information, visit applydigital.com.
 
LOCATION: APPLY is hybrid/remote-friendly. The preferred candidate should be based in Latin America, preferably working in hours that align to PT (Pacific Timezone) or ET (Eastern Timezone). Candidates located in Santiago, Chile are able to work out of our Santiago office as remote/hybrid employees. Candidates located outside of Santiago, Chile will be fully remote employees.
 
THE ROLE:

Apply Digital  is looking for an SRE Engineer to join our globally distributed team. This is a hybrid SRE/Service Desk role designed for someone who is passionate about reliability engineering and comfortable supporting day-to-day operational needs across multiple UK e-commerce clients.

You will be a key contributor in maintaining the health and performance of client platforms, responding to incidents, and continuously improving observability and operational processes. While your primary focus is SRE, you will collaborate closely with the Service Desk team to support triaging, escalation, and resolution workflows.

This role is ideal for someone with 2–3 years of experience who thrives in a fast-paced, multi-client environment, values clear documentation, and is comfortable working with a high degree of autonomy during their shift.

WHAT YOU’LL DO

  • Monitor platform health across multiple client environments using tools like Grafana and Prometheus, or other monitoring tools

  • Respond to and triage incidents, following established runbooks and escalation paths

  • Participate in post-incident reviews and contribute to postmortem documentation

  • Support the Service Desk team with technical triaging, incident classification, and resolution

  • Maintain and improve observability dashboards, alerts, and SLI, and SLO tracking

  • Write and maintain runbooks, operational documentation, and knowledge base articles

  • Identify recurring issues and propose automation or process improvements to reduce toil

  • Participate in on-call rotation covering weekends (alternating schedule — one weekend on, one weekend off)

  • Collaborate with the EMEA team during shift overlap to ensure smooth handoffs and continuity

  • Support root cause analysis and contribute to continuous improvement initiatives

WHAT WE’RE LOOKING FOR:

  • Strong proficiency in English (written and verbal communication) is required

  • 2–3 years of experience in SRE, platform operations, or a technical Service Desk role

  • Experience with monitoring and observability tools such as Grafana, Prometheus, or equivalent

  • Solid understanding of incident management processes (triaging, escalation, postmortems)

  • Experience supporting e-commerce platforms

  • Scripting skills in shell and/or Python for automation and operational tasks

  • Familiarity with containerization concepts (Docker, Kubernetes) at an operational level

  • Experience working in Agile environments and using ticketing tools (e.g. Jira)

  • Comfort working independently during early-morning shifts with minimal supervision

  • Strong documentation habits and attention to detail

  • Experience with Agile processes, testing, and code review

  • Strong experience with scripting - shell, Python, etc.

  • Excellent customer service attitude, communication skills (written and verbal), and interpersonal skills

  • Excellent analytical and problem-solving skills

  • Ability to communicate effectively with technical and non-technical stakeholders. You should feel comfortable explaining technical concepts in simple terms

  • Experience working in fast-paced, Agile environments, balancing priorities across multiple projects

  • NICE TO HAVE: 

  • Experience with Google Cloud Platform (GCP) or other major cloud providers (AWS, Azure)

  • Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI)

  • Basic experience with Infrastructure as Code tools such as Terraform

  • Basic knowledge of AIOps concepts and their application in operational workflows

  • SRE or cloud certifications (Google Cloud, AWS, Kubernetes)

LIFE AT APPLY
 
People are at the core of everything we do at APPLY. We provide you with modern tools, systems and approaches, value your time, safety, and health, and strive to build a work community where you can thrive and grow. Here are a few benefits we offer to support you:
 
Agentic Delivery: Our people work in a modern way to deliver client outcomes. Broaden your skills on a range of engagements with international brands that have a global impact.
An inclusive and safe environment: We’re truly committed to building a culture where you are celebrated and everyone feels welcome and safe.
AI & Strategic Upskilling: Accelerate your professional growth with generous training budgets and mentorship, with a specific focus on Agentic AI expertise and the critical human skills required for the future of work.
Generous vacation policy: Work-life balance is key to our team’s success, so we offer ample time away from work to promote overall well-being.
Flexible work arrangements: We work in a variety of ways, from remote, to in-office, to a blend of both.
 
APPLY is a safe, respectful, and inclusive community where differences are celebrated. We are committed to equal opportunity and fostering a workplace where everyone belongs. Learn more in our Diversity, Equity, and Inclusion (DEI) section. For recruitment accommodations, please email [email protected].

Skills Required

  • Strong proficiency in written and verbal English
  • 2-3 years of experience in SRE, platform operations, or technical Service Desk work
  • Experience with monitoring and observability tools such as Grafana and Prometheus or equivalents
  • Understanding of incident management, triage, escalation, and postmortems
  • Experience supporting e-commerce platforms
  • Strong scripting skills in Shell, Python, or similar languages
  • Operational familiarity with Docker and Kubernetes
  • Experience working in Agile environments and using ticketing tools such as Jira
  • Ability to work independently during early-morning shifts with minimal supervision
  • Strong documentation habits and attention to detail
  • Experience with Agile processes, testing, and code review
  • Excellent customer service, communication, analytical, and problem-solving skills
  • Ability to explain technical concepts to technical and non-technical stakeholders
  • Experience balancing priorities across multiple projects in fast-paced Agile environments
  • Experience with Google Cloud Platform or another major cloud provider
  • Familiarity with CI/CD pipelines such as GitHub Actions or GitLab CI
  • Basic experience with Infrastructure as Code tools such as Terraform
  • Basic knowledge of AIOps concepts
  • SRE or cloud certifications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Vancouver
449 Employees
Year Founded: 2016

What We Do

Digital to our core, we are purpose built to transform possibilities for people. We solve complex problems with well-executed solutions tailor-made for continuous growth — we’re ambitious and our clients are too. We work with well-funded start-ups, global brands, and Fortune 1000 companies spanning industries and audiences, including EA, Moderna, League Health, and Realtor.com.

Similar Jobs

AssureSoft Logo AssureSoft

Site Reliability Engineer

Information Technology • Consulting
Remote
11 Locations
272 Employees

Kraken Digital Asset Exchange Logo Kraken Digital Asset Exchange

Site Reliability Engineer

Blockchain • Financial Services • Cryptocurrency • Web3
Remote
12 Locations
2900 Employees

Prediktive Logo Prediktive

Site Reliability Engineer

Information Technology • Professional Services • Software
Remote
11 Locations
119 Employees

Alpaca Logo Alpaca

Senior Site Reliability Engineer

Fintech • Information Technology
Remote
13 Locations
132 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account