Infrastructure Reliability Engineering Lead

Posted 4 Days Ago
Be an Early Applicant
London, Greater London, England, GBR
In-Office
Senior level
Fintech • Financial Services
The Role
Lead and develop an Infrastructure Reliability Engineering team to implement SRE practices across infrastructure platforms. Drive observability, resilience validation, automation (IaC), operational readiness and SLIs/SLOs, reduce operational toil, and improve service reliability. Collaborate with operations, security and delivery teams to identify risks, run resilience testing, govern observability standards, and deliver engineering improvements for secure, recoverable, and supportable infrastructure services.
Summary Generated by Built In
Infrastructure Reliability Engineering Lead

Shift Pattern:

Standard 40 Hour Week (United Kingdom)

Scheduled Weekly Hours:

40

Corporate Grade:

D - Assistant Vice President

Reporting Line:

(UK Division) Information Technology

Location:

UK-London

Worker Type:

Permanent

The Infrastructure Reliability Engineering Lead is responsible for leading and developing the Infrastructure Reliability Engineering team, ensuring the successful implementation of reliability engineering practices across the infrastructure platforms that underpin LME services. Working under the direction of the Infrastructure Reliability Engineering Senior Manager, the role translates Infrastructure Reliability Engineering strategy, standards and governance into practical engineering outcomes and measurable improvements in service reliability, resilience and operational risk reduction.

The role is accountable for driving engineering excellence across observability, resilience engineering, automation, operational readiness and service reliability. The postholder will lead a team of specialist Infrastructure Reliability Engineers, ensuring infrastructure services are designed, operated and continuously improved in a secure, recoverable, observable and supportable manner.

The Infrastructure Reliability Engineering Lead will work closely with Compute Service Operations, Information Security, Application Delivery and other technology teams to identify reliability risks, improve platform resilience, reduce operational toil through automation and ensure infrastructure services meet business, regulatory and operational requirements.

The initial emphasis of the role will be to establish and mature Infrastructure Reliability Engineering practices, operational readiness assessments, resilience validation, platform observability standards and reliability-focused automation across critical infrastructure services.
 

LME: 

LME Group is the world centre for industrial metals trading and clearing. Most of the world’s non-ferrous metals business is conducted in the LME. 

The metals community uses the LME, a member of HKEX Group, as a venue to transfer or take on price risk, as a physical market of last resort and as the provider of transparent global reference prices. 

 Overall Purpose of Role:

The Infrastructure Reliability Engineering Lead is responsible for leading and developing the Infrastructure Reliability Engineering team, ensuring the successful implementation of reliability engineering practices across the infrastructure platforms that underpin LME services. Working under the direction of the Infrastructure Reliability Engineering Senior Manager, the role translates Infrastructure Reliability Engineering strategy, standards and governance into practical engineering outcomes and measurable improvements in service reliability, resilience and operational risk reduction.

The role is accountable for driving engineering excellence across observability, resilience engineering, automation, operational readiness and service reliability. The postholder will lead a team of specialist Infrastructure Reliability Engineers, ensuring infrastructure services are designed, operated and continuously improved in a secure, recoverable, observable and supportable manner.

The Infrastructure Reliability Engineering Lead will work closely with Compute Service Operations, Information Security, Application Delivery and other technology teams to identify reliability risks, improve platform resilience, reduce operational toil through automation and ensure infrastructure services meet business, regulatory and operational requirements.

The initial emphasis of the role will be to establish and mature Infrastructure Reliability Engineering practices, operational readiness assessments, resilience validation, platform observability standards and reliability-focused automation across critical infrastructure services.

Key responsibilities:

Team Leadership & Management

  • Lead, coach and develop the Infrastructure Reliability Engineering team, setting clear direction, priorities, objectives and performance expectations.
  • Build and maintain a high-performing engineering team through effective recruitment, performance management, coaching, mentoring and succession planning.
  • Allocate engineering resources effectively across reliability improvement initiatives, resilience programmes, operational readiness activities and platform engineering priorities.
  • Foster a culture of engineering excellence, accountability, continuous improvement and shared ownership of service reliability.

Reliability Engineering

  • Lead implementation and continual improvement of Infrastructure Reliability Engineering practices across infrastructure and platform services.
  • Define, measure and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability metrics and reliability measures across critical infrastructure services.
  • Identify, assess and drive remediation of reliability risks and systemic weaknesses before they impact production services.
  • Lead reliability reviews, trend analysis and engineering improvement initiatives focused on increasing resilience and reducing operational risk.

Observability & Service Health

  • Establish and govern observability standards across infrastructure platforms including monitoring, telemetry, logging, tracing and alerting.
  • Ensure meaningful operational visibility and actionable service health monitoring across infrastructure services.
  • Lead improvements in monitoring quality, alert effectiveness and operational insights through data-driven analysis.
  • Ensure infrastructure teams have the observability capabilities required to support effective operational management and rapid issue diagnosis.

Resilience Engineering & Operational Readiness

  • Lead resilience validation activities including failover testing, disaster recovery testing, recovery assurance exercises and scenario-based resilience assessments.
  • Establish and maintain Operational Readiness standards, ensuring services meet agreed supportability, recoverability, observability and operational acceptance criteria before entering production.
  • Ensure recovery capabilities are regularly tested and aligned to agreed business recovery objectives.
  • Drive engineering improvements that enhance service continuity, fault tolerance and recovery capability across critical infrastructure services.

Automation and Platform Improvement

  • Lead adoption of Infrastructure as Code, configuration management and automation practices across infrastructure services.
  • Drive standardisation, repeatability and reduction of operational toil through automation and engineering improvements.
  • Support the continual improvement of platform reliability, scalability, recoverability and operational efficiency through engineering-led initiatives.
  • Promote best practice engineering approaches across cloud-native, virtualised and enterprise infrastructure platforms.

Governance, Risk & Stakeholder Management

  • Support the Infrastructure Reliability Engineering Senior Manager in implementing and maturing the Infrastructure Reliability Engineering operating model, standards and governance frameworks.
  • Provide reliability engineering expertise and technical leadership across projects, platform initiatives and service improvement programmes.
  • Work collaboratively with Information Security, Service Operations and Delivery teams to embed reliability, resilience and operational readiness into platform and service design.
  • Manage stakeholder relationships and provide reporting, insights and recommendations regarding service reliability, resilience and operational risk.

PERSON SPECIFICATION:

Academic and Professional Qualifications Required:

  • Bachelor's degree in Computer Science, Engineering, Information Technology or a related technical discipline.
  • Significant experience in Infrastructure Reliability Engineering, Site Reliability Engineering, Platform Engineering or Infrastructure Engineering disciplines.
  • Demonstrable experience leading technical engineering teams within complex enterprise environments.

Professional certifications or equivalent practical experience related to Linux, cloud-native platforms, observability, automation, infrastructure engineering or reliability engineering would be advantageous.

Required Knowledge and Level of Experience:

Candidates should have significant experience operating within complex enterprise technology environments, ideally within financial services or other highly regulated organisations, together with a strong track record of leading teams responsible for critical infrastructure services and engineering outcomes.

The successful candidate will demonstrate strong people leadership, sound technical judgement and the ability to influence engineering decisions across multiple technology domains. They will possess broad expertise in reliability engineering, resilience, observability, automation and platform engineering, while remaining capable of contributing hands-on where required.

Candidates should possess extensive experience improving infrastructure reliability through engineering practices rather than operational intervention, including the use of automation, observability, resilience testing and service measurement.

The successful candidate must be capable of providing technical leadership across specialist engineering disciplines while building team capability and supporting strategic direction established by the Infrastructure Reliability Engineering Senior Manager.

Technical Skills set and Core Competencies Required for Role:

Candidates should demonstrate strong technical breadth across Infrastructure Reliability Engineering disciplines and possess sufficient depth to provide technical leadership, challenge engineering decisions and guide implementation approaches across specialist engineering teams:

  • Strong understanding of Infrastructure Reliability Engineering and Site Reliability Engineering principles and practices.
  • Experience defining and improving SLIs, SLOs, availability targets and reliability metrics.
  • Strong scripting/automation capability (e.g., Python, Bash, Ansible, Terraform, GitOps).
  • Experience with CI/CD pipelines (GitHub Actions, Jenkins, Azure DevOps, GitLab, etc).
  • Strong observability skills—including metrics, logs, distributed tracing, and tools such as Prometheus, Grafana, ELK, OpenTelemetry, Jaeger.
  • Good understanding of enterprise infrastructure platforms including Linux, virtualisation, container platforms and cloud-native technologies.
  • Strong understanding of resilience engineering, disaster recovery, and service continuity principles.
  • Experience conducting root cause analysis and driving long-term reliability improvements.
  • Understanding of operational readiness, service transition and supportability principles.
  • Knowledge of infrastructure security, hardening, compliance controls and operational risk management.
  • Understanding of ITIL-aligned Incident, Problem, Change and Service Management disciplines.

The following experience would be advantageous:

  • Experience establishing or maturing Infrastructure Reliability Engineering, Site Reliability Engineering or Platform Engineering capabilities.
  • Experience operating within financial services, exchanges, clearing organisations or other highly regulated environments.
  • Experience supporting Kubernetes, OpenShift or cloud-native platform ecosystems.
  • Experience implementing enterprise observability solutions and telemetry platforms.
  • Experience leading operational readiness reviews and infrastructure service acceptance activities.
  • Experience delivering large-scale infrastructure automation and standardisation initiatives.
  • Experience working with third-party technology providers, managed service partners and infrastructure suppliers.
  • Experience supporting significant infrastructure transformation or platform modernisation programmes.

Skills and competencies:

  • Strong people leadership, coaching and team development skills.
  • Strong analytical, troubleshooting and problem-solving capability.
  • Excellent stakeholder management and relationship-building skills.
  • Ability to provide technical leadership while maintaining focus on delivery, governance and business outcomes.
  • Strong planning, prioritisation and organisational skills.
  • Excellent written, verbal and presentation skills.
  • Strong focus on automation, standardisation and continuous improvement.
  • Data-driven and evidence-based approach to decision-making.
  • Demonstrates ownership, accountability and engineering excellence.
  • Comfortable operating within high-pressure, business-critical and regulated environments.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, Information Technology or related discipline.
  • Significant experience in Infrastructure Reliability Engineering, Site Reliability Engineering, Platform Engineering or Infrastructure Engineering.
  • Demonstrable experience leading technical engineering teams within complex enterprise environments.
  • Strong scripting and automation capability (Python, Bash, Ansible, Terraform, GitOps).
  • Experience with CI/CD pipelines (GitHub Actions, Jenkins, Azure DevOps, GitLab).
  • Strong observability skills including metrics, logs, distributed tracing and tools such as Prometheus, Grafana, ELK, OpenTelemetry, Jaeger.
  • Good understanding of enterprise infrastructure platforms including Linux, virtualization, container platforms and cloud-native technologies.
  • Strong understanding of resilience engineering, disaster recovery and service continuity principles.
  • Experience conducting root cause analysis and driving long-term reliability improvements.
  • Understanding of operational readiness, service transition and supportability principles.
  • Knowledge of infrastructure security, hardening, compliance controls and operational risk management.
  • Understanding of ITIL-aligned Incident, Problem, Change and Service Management disciplines.
  • Professional certifications or equivalent practical experience related to Linux, cloud-native platforms, observability, automation or reliability engineering.
  • Experience supporting Kubernetes, OpenShift or cloud-native platform ecosystems.
  • Experience establishing or maturing Infrastructure Reliability Engineering, Site Reliability Engineering or Platform Engineering capabilities.
  • Experience operating within financial services, exchanges, clearing organisations or other highly regulated environments.
  • Experience delivering large-scale infrastructure automation and standardisation initiatives.
  • Experience working with third-party technology providers, managed service partners and infrastructure suppliers.
  • Experience supporting significant infrastructure transformation or platform modernisation programmes.

Hong Kong Exchanges Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Hong Kong Exchanges and has not been reviewed or approved by Hong Kong Exchanges.

  • Retirement Support Employer retirement contributions and provident fund structures are described as notably above statutory baselines, with certain entities in the group offering even higher employer pension rates. This strengthens perceived long-term value and helps total compensation feel more robust.
  • Healthcare Strength Core coverage includes medical and dental insurance alongside life and personal-accident protection, with health checkups and comprehensive plans highlighted. This breadth of health protection is seen as a meaningful pillar of the package.
  • Leave & Time Off Breadth Paid leave spans multiple categories, including parental and volunteering time, in addition to standard annual and sick leave. This variety adds non-cash value and supports work-life needs.

Hong Kong Exchanges Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Central
1,723 Employees
Year Founded: 2000

What We Do

HKEX Group is a global exchange group, operating dynamic and integrated financial markets in Asia and Europe. From our home in the financial hub of Hong Kong and an additional base in London, we provide world-class facilities for trading and clearing securities and derivatives in Equities, Commodities, Fixed Income and Currency. Uniquely positioned at the intersection of Chinese and international capital flows, Hong Kong has long been Connecting China with the World. With the accelerated opening-up of China’s capital markets, HKEX continues to be at the forefront of this historic transition, which we believe will Shape the Global Market Landscape

Similar Jobs

Hybrid
London, Greater London, England, GBR
289097 Employees

Mastercard Logo Mastercard

Senior Specialist, Marketing

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
London, Greater London, England, GBR
38800 Employees

Navan Logo Navan

Chief Operating Officer

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
London, Greater London, England, GBR
3300 Employees

Boeing Logo Boeing

Equipment Maintenance Specialist

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Sheffield, South Yorkshire, England, GBR
170000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account