Reliability Engineer

Reposted 16 Days Ago
Be an Early Applicant
Hiring Remotely in Hong Kong
Remote
Mid level
Fintech • Financial Services
The Role
The Trading Operations Engineer will optimize trading systems, resolve technical issues, manage infrastructure, and collaborate with cross-functional teams to enhance system performance.
Summary Generated by Built In

Flow Traders is hiring Reliability Engineers to safeguard the production performance of our global trading platform and the technology estate around it. Undetected degradation costs P&L by the minute.

Most of what makes failures smaller, shorter and easier to contain happens before anything breaks. You automate the repetitive parts of triage, keep alerts to the ones that need a human, and close monitoring gaps. Root causes go back to the teams that own them.

When something does break, control of the room is yours. You're first on it: you size it up, decide who is on the call, and run the response as Incident Commander under our Global Incident Management framework. You coordinate and delegate rather than dropping into the debugging, and your decisions hold even when the room is more senior than you.

You'll sit alongside traders, developers, infrastructure and specialist teams, with the standing to push any of them when operational standards slip.

What You Will Do

Making failures smaller

  • Automate repeatable triage work so first-line responds faster and more consistently, including alert enrichment, routing, correlation and operational workflows

  • Push for monitoring, alerting and visibility where gaps exist

  • Track reliability and availability across critical trading applications and the platforms they depend on, and work with users, development teams and IT to find where service levels are degrading

  • Call out weak ownership, missed SLAs, poor alerts and ineffective runbooks, and drive the owning teams to fix them

Running the response

  • Triage incoming alerts, issues and escalations, and assess impact, urgency and ownership

  • Decide when incident criteria are met, declare the incident, and act as Incident Commander

  • Coordinate responders and stakeholders, and keep incident calls focused on facts, mitigation and recovery

  • Maintain clear timelines, actions and status updates throughout an incident

  • Recover or stabilise systems using approved runbooks, and escalate cleanly through the defined support and development path when the issue goes beyond documented recovery steps

  • Perform common operational tasks across adjacent teams where needed

After the incident

  • Support PIR follow-up and recurring issue review

  • Hand over cleanly between EMEA, AMER and APAC under one global model, one incident standard, one handover process


What You Will Need to Succeed

The role

  • Experience in production operations, SRE, NOC/command centre, trading operations or a similar first-line technical role, ideally in a trading, financial services or other latency-sensitive environment

  • Track record of running or coordinating major incidents, and comfort taking command of a call with senior people on it

  • Strong triage and prioritisation. You can separate facts from assumptions under time pressure and keep the response moving

  • Clear verbal and written communication. Your status updates are readable by a trader and an engineer at the same time

  • Strong judgment and escalation discipline. You know when to keep going and when to pull in a specialist

  • Willingness to hold the line on process, and to push back when poor operational behaviour creates risk for trading

The technology

  • Technically broad rather than deep. You need enough understanding of how most teams operate to be useful across domains, not to be the specialist resolver

  • Solid Linux and networking fundamentals, and the ability to read alerts, logs, dashboards and symptoms quickly

  • Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data)

  • Familiarity with incident and observability tooling: PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search

  • Scripting and automation ability (Python preferred; Bash, Go a plus) applied to triage, enrichment, routing and correlation rather than to product code

  • Exposure to containerised and cloud-hosted production systems (Kubernetes, Docker, GCP) is a plus


What Good Looks Like

Six months in, a strong Reliability Engineer here has already made the next incident smaller. They've automated something that used to be manual, retired alerts nobody could act on, and closed a monitoring gap that was costing us detection time. When something does break they are calm under pressure, clear in communication and disciplined in process. They can control a noisy incident without trying to become the specialist resolver, and they know how to use a runbook safely and when to escalate. And they don't let poor operational standards slide when trading is exposed.

 

What We Offer

We like to think that talent grows at Flow and stays at Flow. To ensure this, we provide our employees with an extensive onboarding program, access to Flow Academy, the best working environment, the latest technology and continuous support. We go out of our way to retain the small business feeling with which we started and stimulate innovation and collaboration through teamwork and our non-hierarchical approach. We offer competitive salary, annual discretionary bonus and other fantastic perks and benefits, such as:

  • Flow Academy for continuous learning and opportunities to attend domain-related conferences
  • Comprehensive health insurance coverage
  • In-house lounge with a bar, pool table and console games
  • Daily catered breakfast and lunch with healthy snacks and drinks available throughout the day
  • In-house hairdresser and massage therapist
  • Personal trainers, weekly boot camps and a subsidized gym membership
  • Annual company trip and a variety of events throughout the year
  • Global rotations between our offices worldwide
  • and more!

Flow Traders does not accept unsolicited resumes from any professional staffing or search firms. All resumes, and any other information identifying potential candidates, submitted to any employee at Flow Traders via-email, the Internet or directly without a valid and signed search agreement will be deemed free to contact by Flow Traders without any restrictions and no placement fee of any kind will be paid in the event the candidate is hired by Flow Traders.

Skills Required

  • 3-5 years of experience in Trading Operations, DevOps, or Infrastructure Engineering
  • Deep expertise in Unix/Linux systems
  • Strong programming skills in Python
  • Knowledge of TCP/IP and network troubleshooting
  • Experience with Kubernetes, Helm, Docker
  • Hands-on experience with Kafka
  • Familiarity with Terraform and Ansible
  • Strong troubleshooting skills in high-performance systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Amsterdam
626 Employees
Year Founded: 2004

What We Do

Since 2004, Flow Traders has been a principal trading firm and one of the world’s largest liquidity providers, specialised in Exchange Traded Products (ETPs). Our headquarters in Amsterdam and offices in New York, Singapore, Hong Kong, Milan and Cluj accommodate more than 500 employees. Throughout the years, Flow Traders has received the industry’s recognition, winning numerous awards and has been listed as a company since 2015. Our non-hierarchical approach stimulates innovation and development. We value creative minds and out-of-the-box thinkers and challenge them to make full use of their capacities. Our demanding, high-paced environment continuously puts us to the test. Flow Traders fosters a strong team-oriented culture which rewards people for their contributions to the company as a whole rather than only in their direct area of responsibility. At Flow Traders, we focus on professional and personal development; we encourage our people to be the best at anything they do by offering the right training. On top of that, we have our private gym where our health coaches provide our people with the right personal sports programs. We invest in our hard working people, because they hold the key to our success.

Similar Jobs

Remote
Hong Kong
6435 Employees
Remote
Hong Kong
774 Employees

Hyphen Connect Limited Logo Hyphen Connect Limited

Site Reliability Engineer

Agency • Artificial Intelligence • Blockchain • Web3
Remote
Hong Kong
7 Employees

Citadel Logo Citadel

Site Reliability Engineer

Information Technology • Software • Financial Services • Big Data Analytics
In-Office or Remote
4 Locations
4000 Employees
105K-300K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account