Site Reliability Engineer

Posted 21 Days Ago
Be an Early Applicant
27 Locations
Remote
Mid level
Digital Media • Fintech • Gaming • Sports
The Role
Manage and improve cloud infrastructure and Kubernetes platforms (EKS) using GitOps. Own on-call rotations, incident response, SLIs/SLOs, observability stacks, alerting, and automation. Enable rapid multi-country deployments, collaborate with security audits, and mentor junior engineers.
Summary Generated by Built In

What you’ll be doing

  • Work with a team of DevOps and DBA professionals
  • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
  • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
  • Own weekday on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
  • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
  • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
  • Take ownership and responsibility for our cloud operation activities
  • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
  • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
  • Mentoring less experienced team members

What you’ll bring

  • 3+ years DevOps / SRE / platform engineering experience
  • Must be based in Europe 
  • Experience independently leading the planning and deployment of a project
  • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
  • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
  • Experience with Infrastructure-as-Code, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
  • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
  • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
  • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
  • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
  • Experience defining SLIs and SLOs and using them to inform reliability work
  • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
  • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
  • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
  • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
  • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous

Our stack

  • Languages: Java / Spring Boot, Node.js, Python, JavaScript
  • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
  • Cache: ElastiCache, Redis, Valkey
  • Messaging: Apache RocketMQ, AutoMQ, Kafka
  • Networking & Proxy: Nginx, Kong, Cilium, eBPF
  • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
  • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
  • CI/CD: Jenkins, GitHub Actions
  • Metrics: Prometheus, Mimir, Grafana, Alertmanager
  • Logs: Loki, Vector
  • Traces: Tempo, OpenTelemetry, Alloy
  • Profiling: Pyroscope
  • RUM: Grafana Faro, OpenTelemetry SDK
  • Infrastructure as Code: Terraform
  • CDN & Edge: Cloudflare, AWS CloudFront
  • AWS CloudWatch

What’s in it for you

  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities

Interview process

  • Remote video screening with our Talent Acquisition Team
  • Online assessment via Hackerrank
  • Remote video interview with 3 x Team Members (45 mins each, not separate days)

If you’re interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.

Skills Required

  • 3+ years DevOps / SRE / platform engineering experience
  • Based in Europe
  • Experience independently leading the planning and deployment of a project
  • Experience with AWS and cloud resource utilisation (EC2, VPC, Lambda, S3, EBS)
  • Strong understanding of Kubernetes and container orchestration (EKS)
  • Familiarity with GitOps practices
  • Experience with ArgoCD and Helm
  • Infrastructure-as-Code experience, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang
  • Experience with Rust
  • Hands-on experience with observability: Prometheus, Loki, Tempo, Pyroscope, OpenTelemetry
  • Experience with Real User Monitoring (Grafana Faro or OpenTelemetry SDK)
  • Proven on-call and incident response experience, post-mortem leadership
  • Ability to design and maintain alert frameworks to reduce noise and prevent alert fatigue
  • Experience defining SLIs and SLOs and using them to prioritise reliability work
  • Solid networking knowledge (TCP/IP, HTTP)
  • Experience designing systems for high HTTP volumes, high availability
  • Strong understanding of caching and CDNs (HTTP cache, Redis, Memcached, CloudFront, Cloudflare)
  • Excellent Linux troubleshooting and OS parameter optimisation skills
  • JVM optimisation experience
  • Familiarity with service mesh concepts (Cilium-based service mesh)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1,024 Employees
Year Founded: 2013

What We Do

Sporty Group is a global consumer internet and technology company specializing in sports media, gaming, social platforms, and fintech. It creates high-scale consumer products, including sports broadcasting via SportyTV and gaming through SportyBet, serving millions of daily active users across more than 10 countries and 3 continents. The company focuses on delivering world-class user experiences at the intersection of tech and live content.

Similar Jobs

P2P.org Logo P2P.org

Site Reliability Engineer

Information Technology
Remote
28 Locations
179 Employees

Nebius Logo Nebius

Site Reliability Engineer

Artificial Intelligence • Information Technology • Consulting
In-Office or Remote
27 Locations
473 Employees

Zencoder Logo Zencoder

Senior Engineer

Artificial Intelligence • Information Technology • Software
In-Office or Remote
28 Locations
25 Employees

PostHog Logo PostHog

Site Reliability Engineer

Software • Analytics
Remote
26 Locations
60 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account