Digital Site Reliability Engineer

Posted 7 Days Ago
Be an Early Applicant
Madrid, Comunidad de Madrid, ESP
Hybrid
Mid level
Events • Sales • Travel • Hospitality
The Role
Ensures the availability, stability, monitoring, and operational performance of brand web and mobile applications. Leads major incident response, troubleshoots L2/L3 outages and service degradation, performs root cause analysis, coordinates technical bridges, communicates with stakeholders, maintains runbooks and documentation, and drives monitoring, alerting, post-mortem, and long-term reliability improvements.
Summary Generated by Built In
Company Description

Radisson Hotel Group is one of the world's largest hotel groups with ten distinctive hotel brands, and more than 1,460 hotels in operation and under development in 95+ countries. The Group’s overarching brand promise is Every Moment Matters with a signature Yes I Can! service ethos.

People are at the core of our business success and future. Our people are true Moment Makers and together we bring the culture, spirit, environment and opportunities that empower you to be your best, every day, everywhere, every time. Together, we make Every Moment Matter.
 

Job Description

Digital Site Reliability Engineer
 
The Digital Site Reliability Engineer is responsible for ensuring the availability, stability and operational performance of Radisson’s brand web and mobile applications.
 
This role focuses on technical operations, proactive platform monitoring, incident response, root cause analysis, stakeholder communication and support process optimization as part of the Site Reliability Engineering team, working closely with Observability and DevOps specialists, internal teams and external partners to maintain reliable digital services and high-quality customer experience across all digital channels.
 
Role purpose:
1. Reliability & monitoring
• Monitor the health, performance, and availability of digital platforms and business-critical customer journeys.
• Perform detailed investigations (RCA) into recurring issues and service degradations, leveraging Observability platforms.
• Proactively identify service degradations before customer impact.
• Maintain operational runbooks, knowledge base articles, and support documentation.
• Improve operational visibility through dashboards, alerts, reporting, and service health indicators.
• Drive continuous improvements in monitoring coverage and alert quality.
 
2. Incident Management
• Act as technical lead during major incidents affecting digital services.
• Troubleshoot application outages or degraded performance issues. 
• Coordinate technical bridges involving internal teams, external vendors, and business stakeholders.
• Communicate service status, risks, and recovery progress to technical and non-technical stakeholders.
• Ensure proper ticket priorization, triage, routing, and escalation across L1/L2/L3 teams.
• Perform root cause analysis and implement corrective actions end-to-end with development teams.
• Document incidents, follow up on post-mortems, and drive long-term fixes.
 

Qualifications

Must-have experience:
• 2–3 years of experience in application support, site reliability, or service operations environments.
• Experience leading or coordinating major incident bridges involving multiple stakeholders.
• Strong communication skills, with the ability to manage technical and business stakeholders during critical situations.
• Strong experience managing and troubleshooting complex L2/L3 incidents.
• Strong understanding of Incident Management, Problem Management, and Root Cause Analysis processes.
• Hands-on experience with Observability and monitoring tools (e.g. Dynatrace, Grafana, New Relic).
• Proven ability to investigate service degradation, performance issues, and availability incidents.
• Exposure to cloud environments (Azure, AWS, GCP).
 
Highly desirable experience:
• Understanding of modern web applications, APIs, microservices, and distributed systems.
• Familiarity with distributed tracing, synthetic monitoring, and digital experience monitoring.
• Knowledge of CI/CD pipelines, DevOps practices, or Agile methodologies.
• Hospitality industry experience.
• Experience working in multidisciplinary, multicultural and geographically distributed teams.

Additional Information

Why Join Radisson Hotel Group?

Live the Magic of Hospitality - Be part of a team that creates exceptional experiences and memorable moments every day. Let your Yes I Can! spirit shine as you bring hospitality to life.

Build a Great Career - No matter your background or experience, we invest in your growth, learning, and career development—helping you reach your full potential.

Experience the Team Spirit - Join a workplace that’s inclusive, fun, and meaningful. We celebrate diversity, support one another and foster a sense of belonging through our Employee Resource Groups and inclusion initiatives.

Lead with Your Ambition - Your ideas, passion and drive matter! We empower you to make a difference—in hospitality, your community and beyond.

Enjoy Global & Local Perks - No matter where you’re located, you’ll enjoy exclusive global benefits - like special hotel rates for you and your loved ones at our hotels worldwide. Plus, you’ll have access to local perks and rewards tailored to your country, making your experience even more rewarding!

Enjoy benefits such as - up to 53% off your stay as a Team Member at over 1,500 Radisson Hotels worldwide
Guaranteed minimum of 30% off for your Friends & Family
Exclusive Discounts on Breakfast, Food & Beverage, Spa and more

Join us in shaping the future of hospitality! If you’re ready to bring your talent, energy, and passion, we’d love to hear from you.

Apply now and let’s make every moment matter.

We welcome applicants from all backgrounds, abilities, and experiences. If you need any adjustments during the application process, please let us know.

Skills Required

  • 2-3 years of experience in application support, site reliability, or service operations environments
  • Experience leading or coordinating major incident bridges involving multiple stakeholders
  • Strong communication skills with technical and business stakeholders during critical situations
  • Strong experience managing and troubleshooting complex L2/L3 incidents
  • Strong understanding of Incident Management, Problem Management, and Root Cause Analysis processes
  • Hands-on experience with observability and monitoring tools such as Dynatrace, Grafana, or New Relic
  • Ability to investigate service degradation, performance issues, and availability incidents
  • Exposure to cloud environments such as Azure, AWS, or GCP
  • Understanding of modern web applications, APIs, microservices, and distributed systems
  • Familiarity with distributed tracing, synthetic monitoring, and digital experience monitoring
  • Knowledge of CI/CD pipelines, DevOps practices, or Agile methodologies
  • Hospitality industry experience
  • Experience working in multidisciplinary, multicultural, and geographically distributed teams
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Brussels
Year Founded: 1994

What We Do

Radisson Hotel Group is a dynamic hotel company that offers opportunities for growth and aims to create memorable moments through exceptional hospitality. They are involved in sales, revenue management, and meeting/events management.

Similar Jobs

Cloudflare Logo Cloudflare

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Spain
4400 Employees

Mastercard Logo Mastercard

Consultant

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Madrid, Comunidad de Madrid, ESP
38800 Employees
Hybrid
Madrid, Comunidad de Madrid, ESP
289097 Employees

Datadog Logo Datadog

Software Engineering Intern

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Madrid, Comunidad de Madrid, ESP
6500 Employees

Similar Companies Hiring

Posh Thumbnail
Events • Social Media • Software
New York, New York
70 Employees
PRIMA Thumbnail
Travel • Software • Marketing Tech • Hospitality • eCommerce
US
15 Employees
Fairly Even Thumbnail
Hardware • Robotics • Sales • Software • Hospitality
New York, NY
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account