We're building a multi-brand iGaming platform that helps operators run their business in regulated markets worldwide. It brings together three main products — Player Account Management (PAM), Sportsbook, and Casino — all built on a modern microservices architecture using Scala, with a React/TypeScript back office.
We're looking for a Senior SRE to join our team and help us build and maintain reliable, observable, and scalable infrastructure. You'll work closely with developers, own reliability practices, and contribute to the team's DevOps culture.
Requirements
Must Have: Technical Skills
- Linux — confident troubleshooting in terminal, logs, processes, networking basics, resource usage.
- Kubernetes / GKE — strong hands-on understanding of workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting.
- GCP — practical experience with cloud infrastructure, especially around GKE, IAM, networking, load balancing, artifact registry, and production diagnostics.
- Docker / Containers — images, registries, runtime debugging, container lifecycle.
- Helm — understands Helm releases, values, deployment state, and rollback.
- GitOps / FluxCD — able to understand and troubleshoot GitOps deployment flow, drift, image automation, and Git-based rollback.
- CI/CD understanding — can investigate failed pipelines and deployment issues; does not need to be the main pipeline builder if DevOps owns that.
- Observability — Prometheus, Alertmanager, Grafana; understands metrics, logs, traces, dashboards, alerting, and production monitoring.
- Networking fundamentals — DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules.
- SLI / SLO / SLA — can define and apply reliability targets, not just explain the terms.
- Incident response — experience with production incidents, rollback, service recovery, RCA/postmortem.
- Databases / messaging operational basics — PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar; enough to monitor health, migrations, indexes, topics, and failure signals.
- Security / compliance basics — secrets, IAM/RBAC, audit trail, release/change evidence.
Must Have: Reliability / Operational Skills
- Can assess if a release is technically safe to proceed.
- Can define what should be monitored during and after deployment.
- Can verify production health after release.
- Can prepare or validate rollback plans.
- Can lead or support post-release watch.
- Can improve runbooks and incident response procedures.
- Can identify gaps in dashboards, alerts, deployment checks, and release evidence.
- Can work with DevOps without duplicating their ownership.
Must Have: Soft Skills
- Ownership and accountability — follows production issues through to closure.
- Calm under pressure — can operate clearly during incidents, failed deployments, and rollbacks.
- Clear communication — gives concise updates: impact, current action, ETA, risk, decision needed.
- Structured troubleshooting — breaks issues down by app, infra, network, database, config, release diff, external dependency.
- Risk awareness — can say when a release is risky and explain why.
- Collaboration — works well with Dev, QA, Product, Support, Release Manager, and DevOps.
- Documentation discipline — maintains runbooks, rollback steps, incident timelines, and operational evidence.
- Blameless mindset — focuses on root cause and prevention.
- Proactivity — finds missing alerts, dashboards, runbooks, and process gaps before incidents.
- Prioritization — separates real production impact from alert noise.
Nice To Have
- Terraform / IaC experience.
- Advanced GCP infrastructure design.
- ArgoCD or other GitOps tools.
- Advanced database or performance tuning.
- Service mesh experience.
- On-call rotation experience with PagerDuty/Opsgenie or similar.
- Experience facilitating RCA/postmortem sessions.
- Betting/gaming domain experience.
- Good written English for incident updates, release notes, and audit evidence.
- Basic understanding of AI-related concepts: agents, skills, MCP.
- Own production reliability practices together with DevOps and engineering teams.
- Monitor production health and improve observability.
- Support releases, hotfixes, rollback, and post-release watch.
- Validate deployment readiness and rollback readiness.
- Participate in incident response and RCA/postmortem.
- Maintain runbooks, dashboards, alerting rules, and operational documentation.
- Ensure release evidence is traceable and audit-ready.
- Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.
Skills Required
- Strong Linux troubleshooting (terminal, logs, processes, networking, resource usage)
- Kubernetes expertise including workloads, pods, services, ingress/gateway, probes, RBAC, autoscaling, troubleshooting
- GKE and GCP practical experience (IAM, networking, load balancing, artifact registry, production diagnostics)
- Docker and container lifecycle, image and registry management, runtime debugging
- Helm knowledge (releases, values, deployment state, rollback)
- GitOps experience (FluxCD) and ability to troubleshoot GitOps deployment flow and drift
- CI/CD understanding to investigate failed pipelines and deployment issues
- Observability stack experience: Prometheus, Alertmanager, Grafana; metrics, logs, traces, dashboards, alerting
- Networking fundamentals: DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules
- Ability to define and apply SLI/SLO/SLA and monitor against reliability targets
- Incident response experience: handling incidents, rollback, service recovery, RCA/postmortem
- Operational knowledge of databases/messaging: PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK enough to monitor and detect failure signals
- Security and compliance basics: secrets management, IAM/RBAC, audit trails, release/change evidence
- Can assess release safety, define pre/post-deployment monitoring, verify production health, and prepare rollback plans
- Strong ownership, calm under pressure, clear communication, structured troubleshooting, and blameless mindset
- Collaborates effectively with Dev, QA, Product, Support, Release Manager, and DevOps during production events
- Maintains runbooks, dashboards, alert rules, incident timelines, and audit-ready release evidence
- Terraform / Infrastructure as Code experience
- ArgoCD or other GitOps tools experience
- Service mesh and advanced database/performance tuning experience
- On-call rotation experience with PagerDuty/Opsgenie
- Experience facilitating RCA/postmortem sessions and good written English for incident and audit documentation
What We Do
Symphony Solutions is a Dutch-based AI, Cloud, & Agile transformation company. As a premier software provider of custom iGaming, Healthcare, and Airline solutions, it offers expertise in full-cycle software development, cloud engineering, data and analytics, and AI services.








