Senior Site Reliability Engineer (SRE) – OpenShift & Observability

Reposted 5 Days Ago
Be an Early Applicant
Leiden, NLD
In-Office
Senior level
Fintech • Payments • Software • Financial Services
The Role
Operate and evolve an enterprise OpenShift platform, build observability across metrics/logs/traces, define SLIs/SLOs/SLAs, automate via IaC and GitOps, onboard application teams, run incident management and on-call duties to improve platform reliability and scalability.
Summary Generated by Built In

ABOUT US

We’re the world’s leading provider of secure financial messaging services, headquartered in Belgium. We are the way the world moves value – across borders, through cities and overseas. No other organisation can address the scale, precision, pace and trust that this demands, and we’re proud to support the global economy. 

We’re unique too. We were established to find a better way for the global financial community to move value – a reliable, safe and secure approach that the community can trust, completely. We’re always striving to be better and are constantly evolving in an ever-changing landscape, without undermining that trust. Five decades on, our vibrant community reflects the complexity and diversity of the financial ecosystem. We innovate diligently, test exhaustively, then implement fast. In a connected and exciting era, our mission has never been more relevant. Swift now has a presence in 200+ countries and legal territories to serve a community of more than 12,000 banks and financial institutions.   

Role Summary

Join a growing engineering team shaping one of the most strategic platform initiatives at Swift, the world's leading provider of secure financial messaging services. You'll help build, operate and evolve a modern OpenShift platform that will host the next generation of applications across the organisation.

This is a hands-on Senior SRE role focused on platform reliability, observability, automation and operational excellence. You'll work alongside highly skilled engineers across Europe and the US, helping define reliability standards and enabling development teams to successfully adopt cloud-native technologies at enterprise scale.

What you'll do

  • Shape and operate Swift's enterprise OpenShift platform, ensuring high availability, scalability and resilience for critical applications.
  • Lead the design and implementation of observability capabilities across metrics, logs and traces.
  • Establish and evolve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs) that drive measurable reliability improvements.
  • Develop and maintain scalable monitoring and logging solutions using Elasticsearch, Logstash and Kibana (ELK).
  • Partner with application teams to onboard workloads to the platform and promote Site Reliability Engineering best practices.
  • Drive incident management, root cause analysis and post-incident learning to improve platform reliability over time.
  • Design, develop and enhance platform services and capabilities on OpenShift to improve customer experience, security, scalability and operational efficiency.
  • Operate, maintain and continuously improve enterprise OpenShift clusters, including upgrades, lifecycle management, security compliance and platform resilience.
  • Collaborate across two closely connected squads covering platform engineering and application onboarding, providing opportunities to contribute across multiple areas of the platform.
  • Eliminate operational toil through automation and self-service capabilities.
  • Automate operational processes through Infrastructure as Code, GitOps and engineering-first approaches.
  • Participate in on-call rotation, helping ensure the reliability of services supporting business-critical workloads.

What you bring (Must-haves)

  • Proven experience as a Site Reliability Engineer, Platform Engineer or Senior DevOps Engineer in complex production environments.
  • Strong hands-on expertise with OpenShift and/or Kubernetes.
  • Experience designing and operating observability solutions, including logging, monitoring and alerting platforms.
  • Practical experience defining and managing SLIs, SLOs, SLAs and error budgets.
  • Strong understanding of production operations, incident response and reliability engineering principles.
  • A passion for automation, continuous improvement and solving complex technical challenges.

Nice to have

  • Experience with ArgoCD and GitOps operating models.
  • Python development or scripting skills.
  • Infrastructure as Code experience using Terraform, Ansible or similar technologies.
  • Exposure to bare-metal Kubernetes or OpenShift environments.
  • Experience working in large-scale enterprise environments

Hybrid & location

Location: Near Leiden, Netherlands.
Working model: Hybrid working in line with company policy, with regular collaboration across international engineering teams.

What we offer

We give you the freedom to be yourself. We are creating an environment of unique individuals – like you – with different perspectives on the financial industry and the world. A diverse and inclusive environment in which everyone’s voice counts and where you can reach your full potential.

We are committed to an inclusive and accessible recruitment process. If you require a reasonable accommodation related to accessibility during your application or interview, please contact [email protected] or indicate this in your application.

Please note that this mailbox is not monitored for general recruitment enquiries and should only be used for accessibility or accommodation-related requests (for example related to vision, hearing or neurodiversity).

All requests are confidential and will not affect your candidacy.

Don’t meet every single requirement? At Swift, we are dedicated to building a workplace where people can bring their full selves and ideas to the team, so if you are excited about this role, we encourage you to apply even if you do not meet every single qualification.

Skills Required

  • Proven experience as a Site Reliability Engineer, Platform Engineer or Senior DevOps Engineer in complex production environments.
  • Strong hands-on expertise with OpenShift and/or Kubernetes.
  • Experience designing and operating observability solutions including logging, monitoring and alerting platforms.
  • Practical experience defining and managing SLIs, SLOs, SLAs and error budgets.
  • Strong understanding of production operations, incident response and reliability engineering principles.
  • Passion for automation, continuous improvement and solving complex technical challenges.
  • Experience with ArgoCD and GitOps operating models.
  • Python development or scripting skills.
  • Infrastructure as Code experience using Terraform, Ansible or similar technologies.
  • Exposure to bare-metal Kubernetes or OpenShift environments.
  • Experience working in large-scale enterprise environments.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: La Hulpe
4,765 Employees
Year Founded: 1973

What We Do

SWIFT is a global member-owned cooperative and the world’s leading provider of secure financial messaging services. We provide our community with a platform for messaging and standards for communicating, and we offer products and services to facilitate access and integration, identification, analysis and regulatory compliance. Our messaging platform, products and services connect more than 11,000 banking and securities organisations, market infrastructures and corporate customers in more than 200 countries and territories. SWIFT also brings the financial community together – at global, regional and local levels – to shape market practice, define standards and debate issues of mutual interest or concern. For more information, visit www.swift.com or follow us on Twitter: @swiftcommunity

Similar Jobs

UL Solutions Logo UL Solutions

Senior Sales Executive

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Hybrid
8 Locations
15000 Employees

GC AI Logo GC AI

Head of EMEA, GTM

Artificial Intelligence • Legal Tech
In-Office or Remote
28 Locations
130 Employees

Samsara Logo Samsara

Account Development Representative

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Netherlands
4000 Employees

Cloudflare Logo Cloudflare

Senior Sales Director - Enterprise, Benelux

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
2 Locations
4400 Employees
248K-341K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account