Senior Site Reliability Engineer

Posted An Hour Ago
Be an Early Applicant
Hiring Remotely in Hong Kong
Remote
Senior level
Fintech • Payments • Software • Financial Services
The Role
Lead reliability, availability, scalability, observability, incident response, SLO management, and cloud infrastructure operations for a global payment platform. The role manages SEV1/SEV2 incidents, improves Kubernetes and cloud environments, drives automation and post-incident analysis, and strengthens security and resilience in PCI-DSS-regulated systems. It also includes mentoring engineers and coordinating cross-functional responses across global teams.
Summary Generated by Built In
Job Summary

Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Key Responsibilities
  • Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

RequirementsQualifications & Requirements
  • Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Location & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.
Leadership & Operational Excellence
  • Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
  • Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
  • Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
  • Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.

Benefits

- Competitive Package

- A dynamic and innovative team

- Collaborative, inclusive working environment

Skills Required

  • 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles
  • Strong expertise in AWS, Kubernetes/EKS, Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and observability platforms
  • Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration
  • Experience operating payment, banking, fintech, or other highly regulated production environments
  • Knowledge of SRE principles, including SLO/SLI design, error budgets, alert governance, and toil reduction
  • Based in Hong Kong or Shenzhen
  • Excellent written and spoken English
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
53 Employees
Year Founded: 2018

What We Do

Kody is on a mission to make in-person payment acceptance easy. Today, paying in person presents common problems for businesses, such as high costs, long queues, and limited choice of payment methods. Kody fully integrates the payment ecosystem. This way, businesses can offer customers more control over their payment choices to make transactions quicker and simpler. Founded by a small group of final-year high school students and launched in July 2022, 24-year-old founder Yoyo Chang (CEO) studied at the University of Cambridge & York whilst raising US$10M. Today, Kody's platform is growing to connect millions of end-users with venues all over the world.

Similar Jobs

Hyphen Connect Limited Logo Hyphen Connect Limited

Site Reliability Engineer

Agency • Artificial Intelligence • Blockchain • Web3
Remote
Hong Kong
7 Employees

Airwallex Logo Airwallex

Account Executive

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Remote
HK
2300 Employees

Airwallex Logo Airwallex

Growth Executive, SME & Growth, HK

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Remote
HK
2300 Employees

Airwallex Logo Airwallex

Business Development Representative

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Remote
HK
2300 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account