Site Reliability Engineer

Posted 24 Days Ago
Be an Early Applicant
Manchester, Greater Manchester, England, GBR
Hybrid
Entry level
Digital Media • Gaming • Software • Esports • Automation
We’re a world‑leading betting brand, delivering across a broad range of sports and markets globally.
The Role
Build and maintain resilient systems, automation, operational APIs, observability, and infrastructure-as-code tooling. Diagnose distributed-system incidents from edge to origin, manage monitoring and alerting platforms, configure Cloudflare services, participate in incident response and post-mortems, and improve reliability, performance, and operational consistency. Collaborate across SRE, development, and IT Operations while mentoring colleagues and applying AI tools to increase productivity and root-cause analysis.
Summary Generated by Built In
Company Description

We’re one of the world’s leading online gambling companies, revolutionising the industry since 2000. Founded by Denise Coates CBE, we now employ over 10,000 people and serve over 120 million customers in 26 languages.

We empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. This offers employees a wealth of growth opportunities, giving them the opportunity to make a real impact in the world of online gambling. As a forward-thinking company, we’re breaking new ground in software innovation too, redefining what’s possible for our global worldwide.

Our focus on In-Play betting has solidified our market-leading position, featuring more than 1.38 million In-Play sporting events a year. With over 750 concurrent sporting fixtures at peak and more live sports streamed than anyone else in Europe (750,000), we handle over 6 million HTTP requests daily and process more than 1.5 million bets per hour at peak.

Job Description

As a Site Reliability Engineer, you will shape the stability of the systems behind every click, query and live change.

Our Site Reliability team protects and improves the availability, performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate.

You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve operational control. The role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight.

Working across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across the wider engineering community.

This role is eligible for inclusion in the company’s hybrid work from home policy.

Qualifications

  • Software engineering background with Python, Golang, JavaScript or similar language.
  • Knowledge of modern development practices, including testing, source control and delivery lifecycles.
  • An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management.
  • Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty.
  • Proficiency in shell scripting for automation and system management.
  • Experience with Infrastructure as Code, including Terraform and Ansible.
  • Knowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection and traffic management.
  • Ability to troubleshoot distributed systems across edge, network, platform, application, dependency and origin layers.
  • Experience working in a large-scale, 24/7 enterprise where uptime, performance and stability are critical.
  • Practical experience using LLM platforms and coding assistants safely to improve productivity, quality and root-cause analysis.

Additional Information

  • Develop and maintain resilient tools, operational APIs and automation for effective system management.
  • Use orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.
  • Write and contribute to code, telemetry and instrumentation that improve service reliability and observability.
  • Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.
  • Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
  • Diagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.
  • Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence.
  • Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.
  • Drive initiatives that improve reliability, observability, performance and continuous improvement across teams.
  • Mentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.

By applying to us you are agreeing to share your Personal Data in accordance with our Recruitment Privacy Notice - https://www.bet365careers.com/privacy-policy

At bet365, we're committed to creating an environment where everyone feels welcome, respected and valued. Where all individuals can grow and develop, regardless of their background. We're Never Ordinary, and we're always striving to be better. If you need any adjustments or accommodations to the recruitment process, at either application or interview, please don’t hesitate to reach out.

Skills Required

  • Software engineering background with Python, Golang, JavaScript, or a similar language
  • Knowledge of modern development practices, including testing, source control, and delivery lifecycles
  • Understanding of SRE principles, including SLIs, SLOs, reliability measurement, and incident management
  • Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana, or PagerDuty
  • Proficiency in shell scripting for automation and system management
  • Experience with Infrastructure as Code, including Terraform and Ansible
  • Knowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection, and traffic management
  • Ability to troubleshoot distributed systems across edge, network, platform, application, dependency, and origin layers
  • Experience working in a large-scale, 24/7 enterprise where uptime, performance, and stability are critical
  • Practical experience using LLM platforms and coding assistants safely to improve productivity, quality, and root-cause analysis

What the Team is Saying

Nick
Jack
Jack
Justin
Justin
Jenna
Matt
Austen

bet365 Compensation & Benefits Highlights

What else should readers know about bet365's Compensation & Benefits?

We give you the resources so that you can be your best self and feel great. Your wellbeing is our top priority. Here are just some of our life-enhancing benefits:

A healthy body and mind come first

We’ve partnered with ComPsych, who offer a complete support network, including expert advice and compassionate guidance available 24/7. We also offer various fitness resources to help you stay active and healthy.

Giving you financial peace of mind

Whatever life throws at you, we’re here to support along the way. With comprehensive financial benefits, you can have peace of mind knowing we’ve got you covered.

Full-time employees can earn an extra bonus each year, while we celebrate loyalty with rewards for 5, 10 and 25 years of service.

We also offer a 401(k) plan with a company match, as well as company-paid health insurance that includes medical, dental and vision benefits.

Room for personal growth and career development

We recognise and reward your hard work, helping you to learn, grow and thrive. With training and mentoring programmes readily available, you’ll have plenty of opportunities to upskill and progress. Above all, we support everyone in reaching their full potential.

bet365 Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Denver, Colorado
10,000 Employees
Year Founded: 2000

What We Do

We’re one of the world’s leading online gambling companies, revolutionizing the industry since 2000. Founded by Denise Coates CBE, we now employ over 10,000 people and serve over 120 million customers in 26 languages. We empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. This offers employees a wealth of growth opportunities, giving them the opportunity to make a real impact in the world of online gambling. As a forward-thinking company, we’re breaking new ground in software innovation too, redefining what’s possible for our global worldwide. Our focus on In-Play betting has solidified our market-leading position, featuring more than 1.38 million In-Play sporting events a year. With over 750 concurrent sporting fixtures at peak and more live sports streamed than anyone else in Europe (750,000), we handle over 6 million HTTP requests daily and process more than 1.5 million bets per hour at peak.

Why Work With Us

Innovation thrives at bet365. We empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. With opportunities for growth and collaboration, team members have the chance to make a real impact in the world of online betting and gaming.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

bet365 Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We know how important your life is outside of work, which is why many of our roles offer the flexibility to work from home up to 3 days a week.

Typical time on-site: Not Specified
Company Office Image
HQDenver, CO
Company Office Image
HQStoke-On-Trent
Barueri
Bogota
Company Office Image
Gibraltar
Company Office Image
Manchester
Marlton, NJ
Company Office Image
Sliema
Sofia
Sydney
Darwin
Learn more

Similar Jobs

bet365 Logo bet365

Software Engineer

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Manchester, Greater Manchester, England, GBR
10000 Employees

bet365 Logo bet365

Software Engineer

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Stoke-on-Trent, Staffordshire, England, GBR
10000 Employees

bet365 Logo bet365

Site Reliability Engineer

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Stoke-on-Trent, Staffordshire, England, GBR
10000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account