Senior Site Reliability Engineer, Vulcan (AI Security product)

Posted 8 Days Ago
Be an Early Applicant
Hiring Remotely in UAE
Remote
Senior level
Artificial Intelligence • Insurance • Cybersecurity • Web3
The Role
Owns on-premise and airgapped Kubernetes deployments, upgrades, migrations, troubleshooting, and production infrastructure operations for enterprise and government clients. Manages Kubernetes, etcd, PostgreSQL replication, distributed storage, registries, logging, networking, certificates, and automation. Provides independent on-site incident response, validates changes before production, communicates with stakeholders, writes operational documentation, and travels to secure data centers for deployment and go-live support.
Summary Generated by Built In
Job Overview

We are looking for a hands-on infrastructure engineer to own the deployment, migration and troubleshooting of on-premise Kubernetes environments for enterprise and government clientsincluding airgapped, high-security data center environments where remote access is not possible.  

This is a client-facing, on-site role: you will be the technical authority in the room, responsible for executing complex infrastructure changes correctly the first time, diagnosing failures independently under pressure and communicating clearly with client stakeholders throughout. 

This role carries real ownership; you will be expected to understand the systems deeply enough to make sound judgment calls when things don't go to plan, without waiting on remote support.

In this role, you will play a vital part in supporting our Cybersecurity business, Vulcan. Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security. 

Learn more about us 👉

  • Vulcan product: https://vulcanlab.ai/
  • Vulcan LinkedIn: https://www.linkedin.com/company/vulcanlab-ai/
  • AIFT group: https://aift.io/

Responsibilities
  • Plan and execute on-prem Kubernetes cluster deployments, upgrades and infrastructure migrations (including IP re-addressing, certificate rotation and cluster reconfiguration) in production and airgapped environments 
  • Diagnose and resolve failures independently on-site 
  • Own the full infrastructure stack end-to-end: Kubernetes control plane and data plane, PostgreSQL (primary/replica replication), distributed storage (e.g. SeaweedFS/Ceph/similar), private container registries and centralized logging (ELK or equivalent)
  • Validate deployment tooling (scripts, installers, automation) thoroughly in lab/staging environments before any client-facing execution 
  • Represent the technical work directly to client stakeholders on-site: explain status, failures and remediation plans clearly 
  • Travel to client data centers (including airgapped/restricted-access sites) as required, sometimes on short notice, for deployment and go-live support
  • Write clear, structured runbooks, decision trees and incident reports that others (including less experienced engineers) can follow under pressure 
  • Escalate risks proactively to internal leadership, not just after something has gone wrong 
Requirements

Technical: 

  • 5-6 years of hands-on experience with Kubernetes in production, including at least one on-premise (not purely cloud-managed) deployment 
  • Solid understanding of etcd internals. Quorum, peer membership, failure recovery, not just kubectl-level familiarity 
  • Experience with kubeadm-based cluster bootstrapping and certificate management (SANs, CA rotation, renewal) 
  • Working knowledge of PostgreSQL replication, Linux networking fundamentals (DNS, NTP, firewalls) and container registries (Docker Distribution or similar) 
  • Comfortable working entirely from the Linux command line, writing and debugging bash scripts and reading unfamiliar automation tooling under time pressure 
  • Experience with at least one distributed storage system (SeaweedFS, Ceph, MinIO or similar) is a strong plus 
  • GPU-enabled Kubernetes nodes (NVIDIA device plugin, container toolkit) experience is a plus, not required 

Working style:

  • Demonstrated ability to work independently in high-pressure, high-stakes environments without live support 
  • Strong incident communication. Can explain technical failures to non-technical stakeholders factually and calmly, without over-promising or minimizing 
  • A track record of validating changes in test environments before touching production and the judgment to insist on this even under deadline pressure 
  • Comfortable with travel, including to secure/restricted facilities where personal devices, internet access or remote assistance may not be available

Nice to have: 

  • Prior consulting, systems integration or professional services experience, ideally on enterprise or government accounts 
  • Experience specifically in the GCC/Middle East region, or with government-sector clients 
  • Security background (the ability to reason about access controls, credential handling and airgapped operational discipline is valuable given the environments involved) 
Interview Process
  • HR phone interview: 1 hour
  • Online interview: 1.5 - 2 hours, meet with hiring manager
  • Online interview: 1 hour, meet with hiring team
Why Join Us? 
  • Innovative Environment: Be part of a company at the forefront of technology to provide security in GenAI, with opportunities to work on groundbreaking projects. 
  • Growth Opportunities: Take your career to new heights with our career development programs and growth-focused culture. 
  • Dynamic Team: Join a multi-cultural and dynamic team of dedicated professionals who inspire and support each other. 
  • CompensationCompetitive salary and benefits package, commensurate with experience and performance. 

Skills Required

  • 5–6 years of hands-on production Kubernetes experience
  • Experience with at least one on-premise Kubernetes deployment
  • Solid understanding of etcd internals, including quorum, peer membership, and failure recovery
  • Experience with kubeadm-based cluster bootstrapping and Kubernetes certificate management
  • Working knowledge of PostgreSQL replication
  • Knowledge of Linux networking fundamentals, including DNS, NTP, and firewalls
  • Experience with container registries such as Docker Distribution or similar
  • Ability to work from the Linux command line and write and debug Bash scripts
  • Ability to work independently in high-pressure, high-stakes environments
  • Strong incident communication skills with technical and non-technical stakeholders
  • Ability to validate changes in test environments before production deployment
  • Willingness and ability to travel to secure or restricted client facilities
  • Experience with distributed storage systems such as SeaweedFS, Ceph, or MinIO
  • Experience with GPU-enabled Kubernetes nodes, NVIDIA Device Plugin, or NVIDIA Container Toolkit
  • Prior consulting, systems integration, or professional services experience
  • Experience with enterprise or government accounts
  • Experience in the GCC or Middle East region
  • Security background involving access controls, credential handling, or airgapped operations
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Sha tin, Sha tin
43 Employees
Year Founded: 2023

What We Do

AIFT is a group of innovative businesses with a shared vision to become the security layer of the future, providing cybersecurity and insurance protection for digital mega trends including AI, Web3 and future lifestyles.​ ​ AIFT strategically identifies under-served markets within these mega trends where it can quickly make an impact and lead innovation.​ Our businesses include: - OneInfinity: http://oneinfinity.global - Vulcan: http://vulcanlab.ai - OneDegree: https://www.onedegree.hk/en-us - IXT: https://theixt.com/ - Cymetrics: https://cymetrics.io/en-us/ - OneSavie: https://www.linkedin.com/company/onesavie-lab/

Similar Jobs

Cloudflare Logo Cloudflare

Senior Customer Engineer, Sub-Saharan Africa

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Dubai, ARE
4400 Employees

Mastercard Logo Mastercard

Manager - Total Rewards, Benefits & Wellbeing

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Dubai, ARE
38800 Employees

Ericsson Logo Ericsson

Head of IT Infrastructure & Support

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Remote
Dubai, ARE
88000 Employees

Mastercard Logo Mastercard

VP, Product Management - Corporate Solutions, EEMEA

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Dubai, ARE
38800 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account