Senior Site Reliability Engineer

Posted 15 Days Ago
Be an Early Applicant
London, Greater London, England, GBR
In-Office
Senior level
Fintech • Software • Financial Services
The Role
Senior Site Reliability Engineer responsible for maintaining and improving production infrastructure supporting cloud-based core banking and payments platforms. Duties include designing resilient systems, disaster recovery, backups, redundancy, capacity planning, automation, observability, on-call support, production maintenance, documentation, and reliability-focused product development. The role also involves mentoring engineers, collaborating with product teams and customers, and managing complex fleet operations across SaaS environments.
Summary Generated by Built In

Thought Machine's mission is bold – to properly and permanently rid the world's banks of legacy technology. To achieve this, we have developed the foundations of modern banking through core and payments technology which run natively in the cloud. What we are attempting is hard and means we need great people working together to build great technology.

We have grown rapidly in the past few years – growing our team to more than 550 individuals across offices in London, New York, Singapore, Sydney and our newly established Engineering Hub in Lisbon. We have raised more than £500m in funding and our investors include Molten Ventures, Eurazeo, Intesa Sanpaolo, Temasek, Nyca Partners, JPMorgan Chase Strategic Investments, Standard Chartered Ventures, and more.

We have created a culture that enables our team to produce the best work in the industry while ensuring we have fun along the way. We're regularly cited as having a fantastic workplace culture and have been recognised by Sifted magazine as having one of the highest Glassdoor ratings for a UK fintech company and the industry's most generous employee share package. Named one of the world's most innovative fintechs by Global Finance Magazine, we were also recognised by the Financial Times as one of Europe's fastest-growing companies for two consecutive years—and a UK Best Employer for 2026.

Thought Machine’s Site Reliability Engineers are the guardians of mission-critical systems for the world's most influential financial institutions. As a member of our elite, globally distributed team, you'll be entrusted with running and maintaining the robust production infrastructure that powers our customers' cutting-edge Core Banking and Payments platforms. This is an opportunity to make a tangible impact on the global financial landscape while collaborating with brilliant minds to solve complex engineering challenges.

This role will be part of the Site Reliability Engineering team at Thought Machine HQ in London, tackling the challenges of automating complex fleet management operations, mentoring team members, promoting communities of best practice within engineering as well as designing operational processes that provide effective interfaces between Thought Machine and our SaaS customers.

The SRE team is deeply involved in tackling the technical challenges of executing Thought Machine’s growth ambitions - expect to be working with senior stakeholders in the organisation and with our customers, and working on programmes and initiatives that are critical to the success of the company.

Duties:

  • Supporting the product engineering teams in building highly fault-tolerant, scalable applications by participating in design discussions, engaging in RFCs and code reviews.

  • Executing various department strategies - contributing to the design and scoping work for team members around disaster recovery, backup, redundancy and capacity planning activities.

  • Being part of a global on-call rotation responsible for identifying and fixing bottlenecks in SaaS customer environments.

  • Regular maintenance of production systems that host Vault products.

  • Driving the evolution of our SaaS products by defining and designing features that foster exceptional reliability and an unparalleled user experience.

  • Implementing and regularly testing DR strategies to ensure the highest level of resilience and fault tolerance of the platform.

  • Maintain and promote high-quality written documentation of assets, processes and runbooks that are used by the team in their day-to-day operations,

  • Working with your Manager in growing team members in their technical skills as well as their understanding of Vault Products.

Requirements:

  • You have a track record of delivering high-impact projects with focus on long-term scalability, ensuring that human intervention scales sub-linearly with usage growth.

  • You possess an up-to-date understanding of design patterns relevant to hosting and networking architectures.

  • You proactively champion product development, driven by a desire to build truly exceptional products, not just solve immediate challenges.

  • You’re a high-agency individual who can independently drive projects to completion by effectively scaling your individual output with the appropriate delegation of work to team members.

  • You have a strong background working in either Python, Golang having used one of these programming languages to execute a significantly sized project or initiative.

  • You have experience working with Kubernetes or other container orchestration systems.

  • You have experience with automation/configuration management, e.g. Terraform, Puppet, Chef, Ansible.

  • You have expertise in one or more of the following areas: Database Administration, Networking, Observability Tools (such as Prometheus, Jaeger) or automation infrastructure.

  • You have extensive experience working with either GCP or AWS.

Benefits:

  • Highly competitive salary

  • Pension plan (match up to 5%)

  • Life insurance - three times annual salary

  • Competitive maternity (six months fully paid) and paternity leave (four weeks fully paid)

  • Shared parental leave (matched to our maternity leave for the same point in time)

  • 25 days holiday and bank holidays

  • Flexible working hours

  • Cycle-to-work scheme

  • Electric car scheme

  • Season ticket loan

  • Access to outstanding learning materials and courses

  • Sports and hobby clubs, subsidised by Thought Machine

  • All the latest tech you need

  • Start the day properly with fresh fruit and cereals

  • Huge range of healthy (and not-so-healthy) snacks, smoothies and drinks

  • A talented and experienced team as your colleagues

  • An environment where we encourage learning and progress

  • Two charity days a year

  • Weekly food pop-up

We actively hire candidates who demonstrate technical excellence in their field and welcome people of all ages and backgrounds, providing everyone with equal access to professional development. You are encouraged to apply even if your experience doesn't accurately match the job description. We also encourage applications from those with different abilities, including candidates with ADHD, autism, dyslexia or dyspraxia.

Skills Required

  • Track record of delivering high-impact projects focused on long-term scalability
  • Current understanding of hosting and networking architecture design patterns
  • Ability to champion product development and build exceptional products
  • Ability to independently drive projects to completion and delegate work effectively
  • Strong experience with Python or Golang on a significantly sized project or initiative
  • Experience with Kubernetes or other container orchestration systems
  • Experience with automation or configuration management tools such as Terraform, Puppet, Chef, or Ansible
  • Expertise in database administration, networking, observability tools, or automation infrastructure
  • Extensive experience with GCP or AWS
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
617 Employees
Year Founded: 2014

What We Do

Our team’s mission is a bold one – to create technology that can run the world’s banks according to the best designs and software practices of the modern age. In doing so, we will properly and permanently rid the world’s banks of the problems generated by poor technology running on legacy infrastructure. Our solution to this is Vault Core: a complete core banking platform that is capable of being configured easily to suit the needs of any bank. We have built Vault Core from the ground up as a cloud-native, microservices and API-based platform. Thought Machine has a deep culture of engineering excellence, and our approach has engendered a seismic shift in the banking industry. Thought Machine is looking for highly talented individuals to help grow the company and achieve our ambitious goal. We pride ourselves on having an excellent internal culture, where we strive hard to create the best possible working environment; a healthy mix of great technical work, fast pace, a supportive atmosphere, and of course our irreverent sense of fun

Similar Jobs

In-Office
Gloucester, Gloucestershire, England, GBR
35858 Employees

Cisco Logo Cisco

Senior Site Reliability Engineer

Cloud • Information Technology • Internet of Things • Professional Services • Software
In-Office
London, Greater London, England, GBR
77500 Employees
In-Office
2 Locations
2969 Employees

Adaptive (weareadaptive.com) Logo Adaptive (weareadaptive.com)

Site Reliability Engineer

Fintech • Professional Services • Software • Consulting
In-Office
London, Greater London, England, GBR
170 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account