Senior Software Engineer

Posted 4 Days Ago
Be an Early Applicant
London, Greater London, England, GBR
Hybrid
70K-85K Annually
Senior level
HR Tech • Marketing Tech • Software
The Role
Hands-on operations-focused senior software engineer bridging application support and engineering. Provide L2.5 support for PHP services on EKS with MySQL, build automation and runbooks, tune observability (Datadog, Kibana, Heap), participate in on-call rotation, reduce escalations, and improve MTTR through automation, documentation, and collaboration with engineering and platform teams.
Summary Generated by Built In
Reward Gateway, part of Edenred, is a global leader in benefits and employee engagement. We help businesses attract, engage, and retain top talent through strategic reward, recognition, and well-being solutions.
Guided by our shared missions - ‘Making the World a Better Place to Work’ and ‘Enriching Connections, For Good’ - we’re committed to transforming workplaces and improving people’s daily lives.
Our team embodies entrepreneurial spirit, innovation, and respect. We push boundaries, speak up, and stay human, fostering a culture where imagination thrives.
This role offers a hybrid work model to be present in our London office twice a week.
Your Role in our Mission:
This hands-on role sits at the intersection of operational excellence and engineering craft. You’ll bridge the gap between traditional application support and software engineering by executing scripted remediation, configuration management, feature flag operations, safe, bounded code-level fixes, and runbook automation — all under clearly defined guardrails. The goal is to reduce unnecessary L3 escalations while increasing autonomy, quality, and impact for our Application Operations function. You’ll apply these practices across our AWS environment (EKS), PHP services, and MySQL databases, using Datadog as our observability platform, Kibana for log exploration, and Heap to help quantify and understand customer impact.

What You’ll be Doing:
L2.5 Operations Delivery 
  • Provide high-quality, timely L2.5 support for PHP applications running on EKS with MySQL backends, operating within clear guardrails that include configuration changes, feature flag operations, scripted runbooks, and safe, bounded code-level fixes. 
  • Model a shift-left mindset: resolve more at L2.5, automate more, and escalate less, increasing the percentage of incidents resolved without L3 involvement and improving MTTR. 
  • Participate in a healthy, sustainable on-call rotation with fair schedules, clear escalation paths, and strong post-incident learning practices. 
Engineering Practices Within Operations 
  • Apply engineering discipline to operational work: use version control, code review, and testing standards for scripts, runbooks, and automation tooling you produce. 
  • Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns across workloads, services, and operational scenarios. 
  • Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR. 
Observability and Service Readiness 
  • Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app-level alerts and dashboards to minimize noise and surface actionable signals. 
  • Use Kibana to interrogate logs and correlate events with Datadog signals during investigations; improve log usefulness by feeding back patterns for better parsing and context. 
  • Use Heap to triangulate and quantify customer impact (affected flows, cohorts, and volumes) during incidents and problem investigations; incorporate findings into incident timelines and post-incident reviews. 
  • Participate in service onboarding and operability reviews to ensure new and changed services meet defined supportability standards before production. 
  • Contribute to the Service Catalogue with accurate ownership, SLAs/SLOs, runbooks, and escalation paths for supported services. 
Technical Operations and Incident Participation 
  • Act as a first responder for application incidents at L2.5: triage, diagnose, and remediate within guardrails (e.g., safe config changes, feature flag toggles, rolling restarts, cache purges, scripted data fixes). Support major incidents by providing technical context, structured diagnostics, Datadog/Kibana evidence, Heap impact analysis, and coordinated remediation alongside the incident commander. 
  • Use structured diagnostics before escalating — attach clear evidence, reproducibility steps, and impact assessments to every L3/SRE handoff. 
  • Feed operational findings into Problem Management and contribute to post-incident reviews; capture learning in improved runbooks, alerts, and automation. 
Quality, Process, and Continuous Improvement 
  • Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first-contact resolution, and SLO adherence. 
  • Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile-aligned ways of working. 
  • Maintain high standards of documentation — runbooks, known errors, and operational guides are accurate, accessible, and kept up to date. 
Stakeholder Collaboration 
  • Work closely with the Director of Application Operations, Problem Manager, and PETO peers (Platform, Infrastructure, Data, SRE) to ensure a coherent, joined-up operational approach. 
  • Partner with product-aligned engineering teams to understand application architecture, service dependencies, and failure modes; encode this knowledge into operational capabilities and runbooks. 
Scope and Interfaces (complementary to SRE) 
  • In scope: application-centric remediation under guardrails; automation of known issue patterns; high-quality runbooks; structured diagnostics; service readiness/documentation for PHP services on EKS with MySQL; ownership of app-level dashboards/alerts in Datadog, investigative use of Kibana logs, and customer-impact analysis via Heap. 
Working hours and practices for the team
  • Standard hours are 9am - 6pm, Mon - Fri. 
  • 1 day in every 4 is on call, paid at 1.5x hourly rate 
  • On call hours are 6pm - 9am
  • If your on call day falls on a weekend, 24 hour on call cover is required. 

Experience and Skills You Need in this Role:
Essential technical skills:
  • Proven experience in application support or operations engineering in cloud environments, ideally supporting PHP services running on Kubernetes (EKS) with MySQL backends.
  • Hands-on capability in at least one backend language (PHP preferred; Python or similar also valuable) sufficient to read, diagnose, and write safe operational scripts and minor fixes under guardrails.
  • Practical Kubernetes skills for operations: kubectl/Helm basics, investigating pods/deployments, reading logs/events, understanding readiness/liveness probes, and performing safe rollouts/rollbacks within documented guardrails.
  • MySQL operational fluency: connection and pool issues, slow query detection, query plan basics, common remediation patterns (e.g., indexing recommendations to hand to L3, safe data fixes under runbook guardrails), and understanding of replication/backup implications.
  • Strong experience using Datadog (APM/metrics/traces/dashboards/alerts) for investigation and detection; confident using Kibana for log exploration and correlation; ability to leverage Heap to assess user impact and prioritize remediation.
  • Familiarity with ITSM tooling (e.g., Jira Service Management) and ITIL-aligned incident and problem management processes.
  • Strong communication skills; clear, concise documentation; collaborative approach focused on reducing toil, increasing automation, and raising the quality bar.
Nice to have technical skills (but not essential):
  • Python 
    Experience with feature flag platforms and configuration-as-code within safe operational guardrails.
  • Familiarity with AWS services that commonly interface with PHP/EKS workloads (e.g., CloudWatch, ALB, S3, SQS) and how they surface in Datadog and Kibana.
  • Exposure to service onboarding/operability reviews, SLOs, and contributing to a Service Catalogue.
  • Experience balancing incident response with proactive improvement work in Agile contexts; strong documentation discipline.
Responsibilities and Core Duties:
L2.5 Operations Delivery
  • Provide high-quality, timely L2.5 support for PHP applications running on EKS with MySQL backends, operating within clear guardrails that include configuration changes, feature flag operations, scripted runbooks, and safe, bounded code-level fixes.
  • Model a shift-left mindset: resolve more at L2.5, automate more, and escalate less, increasing the percentage of incidents resolved without L3 involvement and improving MTTR.
  • Participate in a healthy, sustainable on-call rotation with fair schedules, clear escalation paths, and strong post-incident learning practices.

Engineering Practices Within Operations
  • Apply engineering discipline to operational work: use version control, code review, and testing standards for scripts, runbooks, and automation tooling you produce.
  • Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns across workloads, services, and operational scenarios.
  • Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR.
Observability and Service Readiness
  • Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app-level alerts and dashboards to minimize noise and surface actionable signals.
  • Use Kibana to interrogate logs and correlate events with Datadog signals during investigations; improve log usefulness by feeding back patterns for better parsing and context.
  • Use Heap to triangulate and quantify customer impact (affected flows, cohorts, and volumes) during incidents and problem investigations; incorporate findings into incident timelines and post-incident reviews.
  • Participate in service onboarding and operability reviews to ensure new and changed services meet defined supportability standards before production.
  • Contribute to the Service Catalogue with accurate ownership, SLAs/SLOs, runbooks, and escalation paths for supported services.

Technical Operations and Incident Participation
  • Act as a first responder for application incidents at L2.5: triage, diagnose, and remediate within guardrails (e.g., safe config changes, feature flag toggles, rolling restarts, cache purges, scripted data fixes). Support major incidents by providing technical context, structured diagnostics, Datadog/Kibana evidence, Heap impact analysis, and coordinated remediation alongside the incident commander.
  • Use structured diagnostics before escalating — attach clear evidence, reproducibility steps, and impact assessments to every L3/SRE handoff.
  • Feed operational findings into Problem Management and contribute to post-incident reviews; capture learning in improved runbooks, alerts, and automation.
Quality, Process, and Continuous Improvement
  • Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first-contact resolution, and SLO adherence.
  • Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile-aligned ways of working.
  • Maintain high standards of documentation — runbooks, known errors, and operational guides are accurate, accessible, and kept up to date.
Stakeholder Collaboration
  • Work closely with the Director of Application Operations, Problem Manager, and PETO peers (Platform, Infrastructure, Data, SRE) to ensure a coherent, joined-up operational approach.
  • Partner with product-aligned engineering teams to understand application architecture, service dependencies, and failure modes; encode this knowledge into operational capabilities and runbooks.
Scope and Interfaces (complementary to SRE)
  • In scope: application-centric remediation under guardrails; automation of known issue patterns; high-quality runbooks; structured diagnostics; service readiness/documentation for PHP services on EKS with MySQL; ownership of app-level dashboards/alerts in Datadog, investigative use of Kibana logs, and customer-impact analysis via Heap.
What Success Looks Like
  • Increased percentage of tickets resolved at L2/L2.5, with reduced unnecessary L3 escalations.
  • A maintained and actively used library of runbooks and automation scripts covering EKS /PHP / MySQL operational scenarios.
  • Measurable reduction in MTTR driven by improved tooling, documentation, and automation.
  • Earned trust of engineering and product peers as a technically credible, collaborative operations engineer.

The Interview Process:
  • Screening call with a member of the Talent Acquisition Team
  • First stage interview with Application Operations Leadership and peer (practical scenario or technical assessment relevant to the L2.5 operating model)
  • Final stage interview with Director or VP
At Reward Gateway | Edenred we are committed to ensuring an inclusive and accessible recruitment process for all candidates. If you have any specific requirements or need reasonable adjustments at any stage of the recruitment journey, please let your Talent Acquisition Partner know. Your needs are important to us, and we want to ensure an equitable experience for every candidate. 
Be comfortable. Be you.
At Reward Gateway, we want all our employees to feel comfortable bringing their passion, creativity and individuality to work. We value all cultures, backgrounds, and experiences, as we truly believe that diversity drives innovation. Express yourself, join our community and help us Make the World a Better Place to Work. 

About
Reward Gateway is culture and client driven. We’re obsessed with putting the “Human” in HR and are proud to have been 100% dedicated to HR for over a decade. Since 2007, we’ve been right by the side of the world’s most innovative HR people, giving them beautiful products and tools they can use to attract, engage and retain their people.The world’s most successful companies treat their people differently. They generate stock market returns of twice their peers and they have half the employee turnover. 76% of CEOs recognize that employee engagement is vital to their success but only 24% say they have a highly engaged company. Bridging that engagement gap is what drives us.

Skills Required

  • Proven experience supporting PHP applications running on Kubernetes (EKS) with MySQL backends
  • Ability to read, diagnose, and write safe operational scripts and minor code fixes (PHP preferred)
  • Practical Kubernetes operational skills (kubectl, Helm, investigating pods/deployments, readiness/liveness probes, safe rollouts/rollbacks)
  • MySQL operational fluency (connection/pooling, slow query detection, query plans, replication/backup awareness)
  • Strong experience using Datadog for APM/metrics/traces/dashboards/alerts
  • Confident use of Kibana for log exploration and correlation
  • Ability to leverage Heap to quantify user impact during incidents
  • Familiarity with ITSM tooling (e.g., Jira Service Management) and ITIL-aligned incident/problem management
  • Use of version control, code review, and testing standards for scripts/runbooks
  • Willingness to participate in on-call rotation and work a hybrid model (London office twice weekly) with standard team hours
  • Strong written and verbal communication; clear documentation discipline
  • Experience with Python
  • Experience with feature flag platforms and configuration-as-code within operational guardrails
  • Familiarity with AWS services interfacing with PHP/EKS workloads (CloudWatch, ALB, S3, SQS)
  • Exposure to service onboarding, SLOs, and Service Catalogue contributions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
600 Employees
Year Founded: 2006

Similar Jobs

Wise Logo Wise

Senior Software Engineer

Fintech • Mobile • Payments • Software • Financial Services
Hybrid
London, Greater London, England, GBR
9000 Employees
88K-111K Annually

monday.com Logo monday.com

Senior Software Engineer

Artificial Intelligence • Productivity • Sales • Software
Hybrid
London, Greater London, England, GBR
3048 Employees

bet365 Logo bet365

Senior Software Engineer

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Manchester, Greater Manchester, England, GBR
10000 Employees

bet365 Logo bet365

Senior Software Engineer

Digital Media • Gaming • Software • Esports • Automation
Hybrid
Manchester, Greater Manchester, England, GBR
10000 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account