Software Engineer, SRE

Posted Yesterday
Be an Early Applicant
Stoke-on-Trent, Staffordshire, England, GBR
Hybrid
Entry level
Digital Media • Gaming • Software • Esports • Automation
We’re a world‑leading betting brand, delivering across a broad range of sports and markets globally.
The Role
Enhance system reliability, observability, and performance through software engineering. Build automation, tooling, telemetry, dashboards, and operational APIs using OpenTelemetry and observability platforms. Define and manage SLIs and SLOs, participate in incident resolution and post-mortems, reduce operational toil with infrastructure automation and AI tools, maintain monitoring systems, and promote reliability best practices across software teams.
Summary Generated by Built In
Company Description

We’re one of the world’s leading online gambling companies, revolutionising the industry since 2000. Founded by Denise Coates CBE, we now employ over 10,000 people and serve over 120 million customers in 26 languages.

We empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. This offers employees a wealth of growth opportunities, giving them the opportunity to make a real impact in the world of online gambling. As a forward-thinking company, we’re breaking new ground in software innovation too, redefining what’s possible for our global worldwide.

Our focus on In-Play betting has solidified our market-leading position, featuring more than 1.38 million In-Play sporting events a year. With over 750 concurrent sporting fixtures at peak and more live sports streamed than anyone else in Europe (750,000), we handle over 6 million HTTP requests daily and process more than 1.5 million bets per hour at peak.

Job Description

As a Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices.

You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing for, and progressing the performance and availability of our critical systems.

Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices.

You will leverage AI tools and LLM platforms in your daily work to reduce toil, drive autonomous operations, and optimise system health, while engineering automation and tooling for effective service management.

Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement.

This role is eligible for inclusion in the Company's hybrid working from home policy.

Qualifications

  • Excellent knowledge of programming languages including Python, Golang and JavaScript.
  • Knowledge and experience of modern software development techniques and lifecycles.
  • Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction.
  • Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty.
  • Proficiency in shell scripting for automation and system management tasks.
  •  Experience with Infrastructure as Code (IaC), automation and orchestration tools such as Ansible and Terraform.
  • Prior experience working in a large scale, 24/7 enterprise where system uptime and stability is of paramount importance to the business.
  • An AI-native engineering approach, with hands-on experience using LLM platforms and coding assistants to improve productivity and quality, and the ability to integrate AI-driven telemetry for advanced observability, predictive insights and root-cause analysis.

Additional Information

  • Developing and maintaining tools that facilitate effective management of our systems, ensuring they are operationally efficient and resilient.
  • Working with automation and orchestration platforms to automate manual activity and reduce toil.
  • Writing and contributing to code that enhances the reliability and observability of services, including telemetry, operational APIs and tooling.
  • Building sophisticated dashboards using a range of telemetry data and dash boarding technologies like Grafana, Splunk and New Relic.
  • Actively participating in live incident resolution and post-mortem analysis, providing effective remediation strategies to improve overall system health and prevent future issues.
  • Driving initiatives to enhance system reliability and observability, contributing to a culture of continuous improvement.
  • Maintaining and administering existing monitoring and analytic toolsets.
  • Mentoring colleagues in use of new technologies or practices.
  • Working with IT Operations to provide and support the use of critical tooling that will enable increasing levels of value to the Business.

By applying to us you are agreeing to share your Personal Data in accordance with our Recruitment Privacy Notice - https://www.bet365careers.com/privacy-policy

At bet365, we're committed to creating an environment where everyone feels welcome, respected and valued. Where all individuals can grow and develop, regardless of their background. We're Never Ordinary, and we're always striving to be better. If you need any adjustments or accommodations to the recruitment process, at either application or interview, please don’t hesitate to reach out.

Skills Required

  • Excellent knowledge of Python, Golang, and JavaScript
  • Knowledge and experience of modern software development techniques and lifecycles
  • Excellent knowledge of Site Reliability Engineering principles, including SLIs and SLOs
  • Knowledge and experience with observability tools and practices, including Splunk, New Relic, Grafana, and PagerDuty
  • Proficiency in shell scripting for automation and system management
  • Experience with Infrastructure as Code, automation, and orchestration tools such as Ansible and Terraform
  • Experience working in a large-scale, 24/7 enterprise where uptime and stability are critical
  • Hands-on experience using LLM platforms and coding assistants in an AI-native engineering approach
  • Ability to integrate AI-driven telemetry for observability, predictive insights, and root-cause analysis

What the Team is Saying

Nintal
Jessica
Pablo
Sthefani
Drew
Elise
Alper
Charlotte
Darren
Sthefani
Alper
Pablo
Charlotte
Ben
Darren
Drew
Ashley
Alper
Charlotte
Cameron
Charlotte
Sthefani’s Story
Alper
Charlotte
Alper
Drew
Pablo
Alper
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Denver, Colorado
10,000 Employees
Year Founded: 2000

What We Do

We’re one of the world’s leading online gambling companies, revolutionizing the industry since 2000. Founded by Denise Coates CBE, we now employ over 10,000 people and serve over 120 million customers in 26 languages. We empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. This offers employees a wealth of growth opportunities, giving them the opportunity to make a real impact in the world of online gambling. As a forward-thinking company, we’re breaking new ground in software innovation too, redefining what’s possible for our global worldwide. Our focus on In-Play betting has solidified our market-leading position, featuring more than 1.38 million In-Play sporting events a year. With over 750 concurrent sporting fixtures at peak and more live sports streamed than anyone else in Europe (750,000), we handle over 6 million HTTP requests daily and process more than 1.5 million bets per hour at peak.

Why Work With Us

Innovation thrives at bet365. We empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. With opportunities for growth and collaboration, team members have the chance to make a real impact in the world of online betting and gaming. As a forward-thinking company, we’re breaking ne

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

bet365 Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We know how important your life is outside of work, which is why many of our roles offer the flexibility to work from home up to 3 days a week.

Typical time on-site: Not Specified
Company Office Image
HQDenver, CO
Company Office Image
HQStoke-On-Trent
Sofia
Marlton, NJ
Darwin
Barueri
Company Office Image
Gibraltar
Company Office Image
Manchester
Sydney
Company Office Image
Sliema
Learn more

Similar Jobs

bet365 Logo bet365

Software Engineer

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Manchester, Greater Manchester, England, GBR
10000 Employees

bet365 Logo bet365

Maintenance Engineer, HVAC

Digital Media • Gaming • Software • Esports • Automation
In-Office
Stoke-on-Trent, Staffordshire, England, GBR
10000 Employees

bet365 Logo bet365

Software Tester, PAY

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Manchester, Greater Manchester, England, GBR
10000 Employees

bet365 Logo bet365

Software Tester, PAY

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Stoke-on-Trent, Staffordshire, England, GBR
10000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account