- Platform Core: Design and operate large-scale distributed data systems
- Own the big data compute and storage infrastructure (MaxCompute/ODPS, Hologres, Spark)
- Build and maintain multi-site task orchestration that dynamically selects engines and enforces policy
- Drive reliability and performance improvements across batch and real-time pipelines
- AI Integration: Build the AI-native platform layer
- Develop and expose MCP (Model Context Protocol) tool interfaces so AI agents can interact with platform APIs
- Build the scheduling and cost-optimization agents that auto-tune resource allocation and alert severity
- Instrument platform telemetry to feed AI-driven SLA monitoring and anomaly detection
- Design context retrieval pipelines (RAG / vector search) for SQL code and config knowledge bases
- Tooling & DX: Evolve the developer experience
- Own the internal data development platform — IDE integrations, code review automation, deployment tooling
- Build APIs-first tools (backfill, ingestion automation) designed for future MCP integration
- Collaborate with data warehouse and service teams to define platform contracts
- Ops & Governance: Drive operational excellence
- Establish SLA benchmarks, cost metrics, and latency dashboards as AI optimization targets
- Build automated incident response and root-cause analysis pipelines
- Define and enforce infrastructure policies across multi-cloud environments
- Scheduling Agent: auto-configure task dependencies, engine selection, cost/performance trade-offs, and alert tiers
- Operations Agent: detect pipeline latency, performance degradation, and schema drift; trigger remediation
- Incident Response Agent: trace SLA breaches to root cause, assign accountability, generate post-mortems
- MCP Tool Layer: design and maintain the cross-platform tool interfaces that all agents call into
- 5+ years of experience building large-scale data platforms (Hadoop/Spark/Flink or equivalent)
- Deep expertise in distributed storage and compute systems (MaxCompute, Hologres, ClickHouse, Hive)
- Strong software engineering skills in Java, Scala, or Python; experience with API-first design
- Hands-on experience with task scheduling systems (Airflow, DolphinScheduler, or in-house equivalents)
- Solid understanding of multi-cloud architectures and cost governance
- Familiarity with LLM integration patterns: tool calling, RAG pipelines, context management
- Experience with MCP or similar agent-tool frameworks is a strong plus
- Passion for building systems that make other engineers 10x more productive
- Competitive total compensation package
- L&D programs and education subsidy for employees' growth and development
- Various team building programs and company events
- Wellness and meal allowances
- Comprehensive healthcare schemes for employees and dependants
- More that we love to tell you along the process!
Skills Required
- 3+ years experience in data platform operations, big data engineering, platform engineering, SRE, or cloud infrastructure operations
- Right to work in Singapore and not require OKX visa sponsorship (priority for applicants who do not need sponsorship)
- Experience with either AWS or Alibaba Cloud
- Hands-on experience with at least one big data or cloud data platform (MaxCompute/ODPS, Hologres, Databricks, StarRocks, Flink, Spark, Hive, Presto, or Trino)
- Strong SQL skills for troubleshooting data issues, job failures, and performance problems
- Solid Linux fundamentals
- Scripting experience with Shell, Python, or similar languages
- Understanding of monitoring, alerting, capacity, access control, and resource management
- Strong incident response and troubleshooting skills, with ability to drive recovery under pressure
- Good communication and collaboration skills in a cross-functional environment
OKX Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about OKX and has not been reviewed or approved by OKX.
-
Fair & Transparent Compensation — Pay is considered competitive or above market, especially in engineering, product, and legal roles across major hubs. This positioning is consistently cited as a major attraction for candidates.
-
Healthcare Strength — Role descriptions indicate comprehensive medical, dental, vision, life, and disability coverage, with employer-paid premiums in some cases. Health coverage is highlighted alongside core benefits like PTO and parental leave.
-
Wellbeing & Lifestyle Benefits — Allowances for education and fitness, meal perks and snacks, team-building budgets, and structured learning programs are described across locations. These extras enhance the total rewards package beyond base pay.
OKX Insights
What We Do
Founded in 2017, OKX is one of the world’s leading cryptocurrency spot and derivatives exchanges. OKX innovatively adopted blockchain technology to reshape the financial ecosystem by offering some of the most diverse and sophisticated products, solutions, and trading tools on the market. Trusted by more than 20 million users in over 180 regions globally, OKX strives to provide an engaging platform that empowers every individual to explore the world of crypto. In addition to its world-class DeFi exchange, OKX serves its users with OKX Insights, a research arm that is at the cutting edge of the latest trends in the cryptocurrency industry. With its extensive range of crypto products and services, and unwavering commitment to innovation, OKX’s vision is a world of financial access backed by blockchain and the power of decentralized finance.
.jpeg)





