As Singapore’s longest established bank, we have been dedicated to enabling individuals and businesses to achieve their aspirations since 1932. How? By taking the time to truly understand people. From there, we provide support, services, solutions, and career paths that meet their individual needs and desires.
Today, we’re on a journey of transformation. Leveraging technology and creativity to become a future-ready learning organisation. But for all that change, our strategic ambition is consistently clear and bold, which is to be Asia’s leading financial services partner for a sustainable future.
We invite you to build the bank of the future. Innovate the way we deliver financial services. Work in friendly, supportive teams. Build lasting value in your community. Help people grow their assets, business, and investments. Take your learning as far as you can. Or simply enjoy a vibrant, future-ready career.
Your Opportunity Starts Here.
ROLE
We are seeking a Data Engineer (Manager) to support the development, maintenance, and improvement of data pipelines for enterprise data warehouse and AI knowledge base platforms within a banking environment. Working under the guidance of senior engineers, you will help transform structured and unstructured data into reliable, reusable datasets that support use cases such as risk management, customer engagement, fraud detection, and intelligent automation. This is a hands-on role with strong learning opportunities across cloud, streaming, and AI-enabled data engineering.
This role reports to Executive Director, Data Engineering, Group Data Office.
KEY RESPONSIBILITIES
Batch & Streaming Data Pipeline Development
- Develop and support data pipelines feeding data lake/ data warehouse platforms, such as Cloudera, AWS Redshift, Snowflakes, under senior guidance.
- Support the implementation of data models for reusable analytical datasets and reporting
- Monitor data pipelines against established SLAs and assist in troubleshooting data quality and processing issues
- Develop and maintain batch processing jobs using Spark, SQL, Python, or Java
- Assist with real-time streaming pipelines using Flink or similar tools
- Build and support ingestion pipelines from APIs, GA4, files, and other data sources
- Support messaging and streaming integrations using Pub/Sub or Kafka
- Write clean, tested, and maintainable SQL and Python code for data ingestion and processing
Orchestration & Automation
- Develop and support Airflow DAGs (or similar tools) for scheduled and event-driven workflows
- Follow established orchestration, coding, and operational standards
Cloud Infrastructure & DevOps
- Support data workloads on Cloudera, or AWS, with exposure to Docker, Kubernetes ((or similar tools)), or Cloud Run
- Use and support established CI/CD pipelines
- Develop and support simple REST APIs and backend services using Python and Flask
- Support Redis caching for data services where required
AI Knowledge Base & RAG Support
- Assist in developing vector database pipelines and embedding generation jobs
- Support document processing and chunking workflows for AI knowledge bases and RAG pipelines
- Help test and monitor semantic search and retrieval quality for AI-facing data layers
- undefined
Cross‑functional Collaboration
- Work with senior data engineers, AI teams, and business stakeholders across Risk, Finance, Marketing, and Operations
- Assist in translating business requirements into clear technical tasks
- Participate in code reviews, testing, documentation, and continuous improvement activities
REQUIREMENTS
- Bachelor’s degree in computer science, information systems, engineering, or a related field
- At least 2 years of relevant experience in data engineering, data platforms, software engineering, or related roles.
- Hands-on experience building or supporting data pipelines using SQL and Python in a data warehouse, data lake, or cloud environment.
- Basic knowledge of AI knowledge base concepts, such as vector databases, embeddings, or semantic search, is a plus
- Exposure to RAG components, such as document chunking, embedding generation, or retrieval, is a plus
- Exposure to data layers supporting LLM or AI agent use cases is a plus.
- Experience or interest in banking and financial services is a plus
Technical Stack
- Data Warehouse / Platform: basic experience with Cloudera, Redshift, or similar platforms
- Batch Processing: working knowledge of SQL and Python; exposure to Spark, ETL, or MapReduce
- Streaming: exposure to Flink or another real-time processing engine is good to have
- Orchestration: basic experience with Airflow or equivalent tools
- Cloud & Infrastructure: exposure to GCP or AWS; basic knowledge of Docker, or Cloud Run
- DevOps / DataOps: basic exposure to CI/CD
- Backend & Serving: Python, Flask, and REST APIs;
Additional Preferred Experience
- Eagerness to learn data engineering, real-time streaming, cloud, and low-latency system design
- Exposure to LLM applications, RAG, or AI agent concepts is welcome but not required
- A quality-focused mindset and willingness to treat data as a reusable product
- Good communication, teamwork, problem-solving, and willingness to learn
Competitive base salary. A suite of holistic, flexible benefits to suit every lifestyle. Community initiatives. Industry-leading learning and professional development opportunities. Your wellbeing, growth and aspirations are every bit as cared for as the needs of our customers.
Skills Required
- Bachelor's degree in computer science, information systems, engineering, or a related field
- At least 2 years of relevant experience in data engineering, data platforms, software engineering, or related roles
- Hands-on experience building or supporting data pipelines using SQL and Python in a data warehouse, data lake, or cloud environment
- Basic experience with Cloudera, Redshift, or similar data warehouse platforms
- Working knowledge of SQL and Python, with exposure to Spark, ETL, or MapReduce
- Basic experience with Airflow or equivalent orchestration tools
- Exposure to GCP or AWS and basic knowledge of Docker or Cloud Run
- Basic exposure to CI/CD
- Experience with Python, Flask, and REST APIs
- Basic knowledge of AI knowledge base concepts, including vector databases, embeddings, or semantic search
- Exposure to RAG components, including document chunking, embedding generation, or retrieval
- Exposure to data layers supporting LLM or AI agent use cases
- Experience or interest in banking and financial services
- Exposure to Flink or another real-time processing engine
- Eagerness to learn data engineering, real-time streaming, cloud, and low-latency system design
- Good communication, teamwork, problem-solving, and willingness to learn
What We Do
OCBC is the longest established Singapore bank, formed in 1932 from the merger of three local banks, the oldest of which was founded in 1912. It is now the second largest financial services group in Southeast Asia by assets and one of the world’s most highly-rated banks, with an Aa1 rating from Moody’s. Recognised for its financial strength and stability, OCBC is consistently ranked among the World’s Top 50 Safest Banks by Global Finance and has been named Best Managed Bank in Singapore by The Asian Banker. OCBC and its subsidiaries offer a broad array of commercial banking, specialist financial and wealth management services, ranging from consumer, corporate, investment, private and transaction banking to treasury, insurance, asset management and stockbroking services. OCBC’s key markets are Singapore, Malaysia, Indonesia and Greater China. It has more than 570 branches and representative offices in 19 countries and regions. These include about 300 branches and offices in Indonesia under subsidiary Bank OCBC NISP, and over 90 branches and offices in Mainland China, Hong Kong SAR and Macau SAR under OCBC Wing Hang. OCBC’s private banking services are provided by its wholly-owned subsidiary Bank of Singapore, which operates on a unique open-architecture product platform to source for the best-in-class products to meet its clients’ goals. OCBC's insurance subsidiary, Great Eastern Holdings, is the oldest and most established life insurance group in Singapore and Malaysia. Its asset management subsidiary, Lion Global Investors, is one of the largest private sector asset management companies in Southeast Asia.



.png)





