Senior Site Reliability Engineer (SRE)

Posted 4 Days Ago
Be an Early Applicant
Hiring Remotely in Bangkok, Phra Nakhon, Bangkok, THA
Remote
Senior level
Fintech • Software • Financial Services
The Role
Lead incident response and troubleshoot production issues by diving into logs, metrics, and application code. Debug applications in Java, Golang, and Python; perform Postgres data investigation and fixes. Maintain Kubernetes clusters and cloud infrastructure, improve observability and automation, drive SRE roadmap, and mentor junior engineers to raise reliability and performance.
Summary Generated by Built In

About:

Step forward into the future of technology with ZILO™.

We’re here to redefine what’s possible in technology. While we’re trusted by the global Transfer Agency sector, our technology is truly flexible and designed to transform any business at scale. We’ve created a unified platform that adapts to diverse needs, offering the scalability and reliability legacy systems simply can’t match.

At ZILO™, our DNA is built on Character, Creativity, and Craftsmanship. We face every challenge with integrity, explore new ideas with a curious mind, and set a high standard in every detail.

We are a team of dedicated professionals where everyone, regardless of their role, drives our progress and creates real impact. If you’re ready to shape the future, let’s talk.

Requirements:
We’re looking for a Site Reliability Engineer to join our SRE team - someone who thrives on solving complex production issues, understands how applications behave in the real world, and takes pride in keeping systems reliable and performant.

This is not a platform engineering role. You won’t just be spinning up Kubernetes clusters or building infrastructure — you’ll be deeply involved in understanding our applications, what they do and how they operate, troubleshooting real-world issues, and working directly on improvements that impact our customers every day.
What You’ll Do

  • Incident Response & Troubleshooting: Investigate and resolve incidents raised by clients, diving into logs, metrics, and application code to identify root causes.
  • Application Debugging: Work across our core stack — Java, Golang, and Python — to trace and fix issues affecting reliability or performance.
  • Data Fixes: Perform data investigation and fixes using Postgres.
  • Operational Excellence: Patch and maintain Kubernetes clusters and other production systems.
  • SRE Roadmap: Contribute to the continuous improvement of our observability, reliability, and automation initiatives.
  • Champion a culture of reliability, automation, and continuous improvement
  • Mentor junior engineers and contribute to technical roadmaps and architectural decisions

Requirements
  • 4+ years of experience in Site Reliability Engineering, DevOps, or related infrastructure/systems engineering roles
  • Solid experience with application debugging in at least one of: Java, Golang, or Python.
  • A good grasp of PostgreSQL — enough to run queries, analyse data, and perform safe fixes.
  • Familiarity with Kubernetes and modern cloud platforms (AWS, GCP, or Azure).
  • Understanding of incident management, observability tools (Grafana, Prometheus, etc.)
  • A mindset focused on reliability, quality, and ownership.

Benefits
  • 23 Annual days holiday (Start and Fixed at 23 days)
  • 15 Public Holidays
  • Provident Fund
  • Health insurance (including immediate family)

Skills Required

  • 4+ years in Site Reliability Engineering, DevOps, or related roles
  • Application debugging experience in at least one: Java, Golang, or Python
  • Proficient with PostgreSQL for queries, data analysis, and safe fixes
  • Familiarity with Kubernetes and maintaining clusters in production
  • Experience with modern cloud platforms (AWS, GCP, or Azure)
  • Understanding of incident management and observability tools (Grafana, Prometheus, etc.)
  • Mindset focused on reliability, quality, and ownership
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
London, England
180 Employees
Year Founded: 2020

What We Do

ZILO is focused on transforming global transfer agency to create sustainable value for firms and the customers they serve. To achieve this, we started with a clean technology slate, a design-driven approach, and a commitment to put people first. Zilo's technology enables firms to replace legacy end-of-life systems, many of which were developed 30+ years ago, and slash costs, risk, and user friction along the way. Our founders, leadership, engineering, and product teams are highly experienced with successful track histories of pioneering innovation-driven businesses, products, and services. Our collective goal is to be the market leading solution in global asset and wealth management

Similar Jobs

Gradion Logo Gradion

Site Reliability Engineer

Artificial Intelligence • Big Data • eCommerce • Retail
In-Office or Remote
5 Locations
100 Employees

Gradion Logo Gradion

Site Reliability Engineer

Artificial Intelligence • Big Data • eCommerce • Retail
Remote or Hybrid
5 Locations
100 Employees

DBS Bank Ltd Logo DBS Bank Ltd

Full-stack Engineer

Fintech • Information Technology • Software • Financial Services
In-Office or Remote
17 Locations
41000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account