Site Reliability Engineer (Managed Patching & Platform Automation)

Posted Yesterday
Be an Early Applicant
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
In-Office
Mid level
Fintech • Payments • Software • Financial Services
The Role
Supports and enhances a managed patching platform through enterprise automation, platform integrations, Linux administration, monitoring, reporting, incident response, and compliance controls. Develops Ansible and Python automation, integrates ServiceNow and CI/CD systems, improves service reliability and scalability, investigates technical issues, maintains operational documentation, and collaborates with infrastructure, security, architecture, and engineering teams.
Summary Generated by Built In

ABOUT US

We’re the world’s leading provider of secure financial messaging services, headquartered in Belgium. We are the way the world moves value – across borders, through cities and overseas. No other organisation can address the scale, precision, pace and trust that this demands, and we’re proud to support the global economy. 

We’re unique too. We were established to find a better way for the global financial community to move value – a reliable, safe and secure approach that the community can trust, completely. We’re always striving to be better and are constantly evolving in an ever-changing landscape, without undermining that trust. Five decades on, our vibrant community reflects the complexity and diversity of the financial ecosystem. We innovate diligently, test exhaustively, then implement fast. In a connected and exciting era, our mission has never been more relevant. Swift now has a presence in 200+ countries and legal territories to serve a community of more than 12,000 banks and financial institutions.   

Experience and Qualifications

  • 3 plus years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or related disciplines
  • Proven experience building and operating enterprise-scale automation solutions
  • Strong hands-on experience in infrastructure automation, Linux administration, and system reliability
  • Experience working within large-scale enterprise environments
  • Bachelor's Degree in Computer Science, Engineering, Information Technology, or equivalent practical experience

Key Responsibilities

Support the Managed Patching Service (MPS)

  • Contribute to the enhancement and continuous improvement of the Managed Patching Service (MPS)
  • Take ownership of assigned technical deliverables and service improvements
  • Support the scalability, reliability, performance, and maintainability of the service
  • Help implement engineering solutions that meet enterprise and regulatory requirements
  • Participate in the evolution of MPS toward a platform-driven and self-service operating model

Design and Build Enterprise Automation Solutions

  • Design, develop, and maintain automation workflows using Ansible Automation Platform
  • Develop reusable automation components, scripts, and operational tooling
  • Integrate automation solutions with ServiceNow, inventory systems, CI/CD platforms, and related enterprise services
  • Apply engineering best practices including testing, version control, peer review, and release management
  • Develop automation and integrations using Python and related technologies where required

Platform Integration and Service Engineering

  • Support onboarding, subscription, scheduling, and maintenance window capabilities within MPS
  • Improve service delivery through automation and standardization
  • Contribute to platform enhancements that improve user experience and operational efficiency
  • Assist in building scalable integration patterns across enterprise platforms
  • Support initiatives that reduce manual effort and improve service adoption

Reliability, Observability and Reporting

  • Implement and maintain operational monitoring, logging, and reporting capabilities
  • Contribute to the definition and measurement of service reliability objectives and operational metrics
  • Improve visibility of patching outcomes, compliance status, and service health
  • Support the development of dashboards and reporting solutions for operational and regulatory requirements
  • Identify opportunities to improve service reliability and reduce operational complexity

Incident, Problem and Operational Management

  • Investigate and resolve complex technical issues affecting service availability or performance
  • Participate in incident response, troubleshooting, and service recovery activities
  • Contribute to root cause analysis and corrective actions following incidents
  • Develop and maintain operational documentation, runbooks, and troubleshooting guides
  • Support continuous improvement initiatives to improve service stability and resilience

Compliance and Governance

  • Support compliance with enterprise security, risk, and regulatory requirements
  • Ensure automation workflows maintain appropriate traceability and auditability
  • Contribute to the implementation of governance controls and operational standards
  • Support evidence collection and reporting requirements for audits and compliance reviews
  • Assist in maintaining service documentation and operational records

Platform Operations and Automation Engineering

  • Contribute to infrastructure automation and platform engineering initiatives beyond Managed Patching Service
  • Apply automation and SRE practices to improve operational efficiency and reliability across related platform services
  • Support service transition, operational readiness, and continuous improvement activities
  • Collaborate with engineering teams to identify automation opportunities and operational improvements
  • Participate in shared engineering responsibilities aligned with evolving business priorities

Collaboration and Technical Contribution

  • Collaborate closely with Infrastructure, Security, Architecture, Service Management, and Engineering teams
  • Work with engineers across teams to identify and resolve technical challenges
  • Participate in design discussions, solution reviews, and technical workshops
  • Share knowledge, best practices, and lessons learned with team members
  • Provide guidance and mentoring to less experienced engineers when required
  • Support service adoption by collaborating with stakeholders and platform consumers

Required Skills

  • Strong expertise in Ansible Automation Platform
  • Strong Linux administration and troubleshooting experience (RHEL preferred)
  • Experience developing automation solutions using scripting languages such as Python
  • Experience integrating enterprise platforms such as ServiceNow, CMDBs, monitoring solutions, and CI/CD tools
  • Strong understanding of Site Reliability Engineering principles and operational practices
  • Experience managing and supporting large-scale infrastructure environments
  • Good understanding of automation governance, change management, and operational controls
  • Strong analytical and problem-solving skills
  • Strong communication and collaboration skills

Preferred Skills

  • Experience with CloudBees, Jenkins, GitHub Actions, or similar CI/CD platforms
  • Familiarity with infrastructure as code practices
  • Experience with observability, monitoring, and enterprise reporting solutions
  • Experience with Power BI or similar reporting tools
  • Experience working in regulated or financial services environments
  • Exposure to platform engineering or self-service operational models

What Success Looks Like (6–12 Months)

  • Managed Patching Service operates reliably and efficiently within its defined scope
  • Automation capabilities are enhanced with reduced manual intervention
  • Service onboarding and operational processes become increasingly standardized
  • Operational and compliance reporting is available and trusted by stakeholders
  • Service reliability and operational performance improve through continuous enhancement
  • Strong collaboration is established across engineering and support teams
  • Contributions made to broader infrastructure automation and platform engineering initiatives
  • Knowledge is actively shared within the team, helping raise overall engineering capability

What we offer

We give you the freedom to be yourself. We are creating an environment of unique individuals – like you – with different perspectives on the financial industry and the world. A diverse and inclusive environment in which everyone’s voice counts and where you can reach your full potential.

We are committed to an inclusive and accessible recruitment process. If you require a reasonable accommodation related to accessibility during your application or interview, please contact [email protected] or indicate this in your application.

Please note that this mailbox is not monitored for general recruitment enquiries and should only be used for accessibility or accommodation-related requests (for example related to vision, hearing or neurodiversity).

All requests are confidential and will not affect your candidacy.

Don’t meet every single requirement? At Swift, we are dedicated to building a workplace where people can bring their full selves and ideas to the team, so if you are excited about this role, we encourage you to apply even if you do not meet every single qualification.

Skills Required

  • 3 or more years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or related disciplines
  • Experience building and operating enterprise-scale automation solutions
  • Strong hands-on experience in infrastructure automation, Linux administration, and system reliability
  • Experience working within large-scale enterprise environments
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience
  • Strong expertise in Ansible Automation Platform
  • Strong Linux administration and troubleshooting experience, preferably RHEL
  • Experience developing automation solutions using Python or similar scripting languages
  • Experience integrating ServiceNow, CMDBs, monitoring solutions, CI/CD tools, and related enterprise platforms
  • Strong understanding of Site Reliability Engineering principles and operational practices
  • Experience managing and supporting large-scale infrastructure environments
  • Understanding of automation governance, change management, and operational controls
  • Strong analytical, problem-solving, communication, and collaboration skills
  • Experience with CloudBees, Jenkins, GitHub Actions, or similar CI/CD platforms
  • Familiarity with infrastructure as code practices
  • Experience with observability, monitoring, and enterprise reporting solutions
  • Experience with Power BI or similar reporting tools
  • Experience working in regulated or financial services environments
  • Exposure to platform engineering or self-service operational models
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: La Hulpe
4,765 Employees
Year Founded: 1973

What We Do

SWIFT is a global member-owned cooperative and the world’s leading provider of secure financial messaging services. We provide our community with a platform for messaging and standards for communicating, and we offer products and services to facilitate access and integration, identification, analysis and regulatory compliance. Our messaging platform, products and services connect more than 11,000 banking and securities organisations, market infrastructures and corporate customers in more than 200 countries and territories. SWIFT also brings the financial community together – at global, regional and local levels – to shape market practice, define standards and debate issues of mutual interest or concern. For more information, visit www.swift.com or follow us on Twitter: @swiftcommunity

Similar Jobs

Capco Logo Capco

Project Manager

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Hybrid
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
6000 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Happy Garden, Petaling, Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
16000 Employees

SailPoint Logo SailPoint

Sales Executive

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
Malaysia
2461 Employees

Cloudflare Logo Cloudflare

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Hybrid
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
4400 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account