Senior Data Engineer – Big Data & Cloudera

Posted 5 Days Ago
Be an Early Applicant
Johannesburg, City of Johannesburg, Gauteng, ZAF
In-Office
Senior level
Agency • Information Technology • Professional Services
The Role
Designs, builds, and optimizes Azure-based data platforms, pipelines, ETL/ELT processes, data models, and data products. Uses Microsoft Fabric, Databricks, Azure services, T-SQL, Python, and Spark. Develops CI/CD and infrastructure-as-code deployments, applies governance and data quality practices, monitors platform performance, and collaborates with analysts, product owners, and stakeholders in an Agile environment.
Summary Generated by Built In

We are looking for a skilled Data Engineer to join a high-performing data engineering environment and contribute to the design, development, integration and optimization of enterprise-scale data solutions.

The ideal candidate will have strong hands-on experience across the Cloudera Data Platform (CDP) and the broader Hadoop ecosystem, with proven expertise in building robust ETL pipelines, processing large datasets and supporting data analytics initiatives.

This is an exciting opportunity for a data engineering professional who enjoys working with Big Data technologies, solving complex data challenges and building scalable solutions that enable smarter business decisions.


Key Responsibilities

  • Design, develop and maintain scalable Big Data and ETL data pipelines.
  • Work extensively with the Cloudera Data Platform (CDP) and Hadoop ecosystem.
  • Develop and optimize data processing solutions using Apache Spark and PySpark.
  • Build and manage data ingestion pipelines using Apache NiFi and Sqoop.
  • Work with HDFS, Hive and Impala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineering solutions using Python and Shell scripting.
  • Integrate and process data from enterprise data sources, including Oracle.
  • Develop, maintain and optimize ETL processes to support business and analytical requirements.
  • Monitor data pipelines and scheduled workloads using Control-M.
  • Perform troubleshooting, performance tuning and root-cause analysis across data processing environments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineering processes.
  • Support data quality, data integrity and data availability across enterprise data platforms.
  • Collaborate with Data Analysts, Developers, Architects, Business Analysts and other technology teams.
  • Contribute to the continuous improvement of data engineering standards, processes and platforms.

Requirements
7–8 years of solid hands-on experience as a platform and data engineer (intermediate to senior level).
  • Design, develop and maintain scalable Big Data and ETL data pipelines.
  • Work extensively with the Cloudera Data Platform (CDP) and Hadoop ecosystem.
  • Develop and optimize data processing solutions using Apache Spark and PySpark.
  • Build and manage data ingestion pipelines using Apache NiFi and Sqoop.
  • Work with HDFS, Hive and Impala for large-scale data storage, processing and querying.
  • Develop complex and optimized SQL queries for data extraction, transformation and analysis.
  • Develop data engineering solutions using Python and Shell scripting.
  • Integrate and process data from enterprise data sources, including Oracle.
  • Develop, maintain and optimize ETL processes to support business and analytical requirements.
  • Monitor data pipelines and scheduled workloads using Control-M.
  • Perform troubleshooting, performance tuning and root-cause analysis across data processing environments.
  • Work within Linux/Unix environments to administer, troubleshoot and automate data engineering processes.
  • Support data quality, data integrity and data availability across enterprise data platforms.
  • Collaborate with Data Analysts, Developers, Architects, Business Analysts and other technology teams.
  • Contribute to the continuous improvement of data engineering standards, processes and platforms.


Skills Required

  • 7-8 years of hands-on experience as a platform and data engineer
  • Strong expertise across the Azure data platform, including Microsoft Fabric, Azure Data Factory, Databricks, ADLS Gen2, Azure Synapse Analytics, Event Hubs, and Stream Analytics
  • Experience designing ETL/ELT processes, data models using Kimball and/or Data Vault 2.0, and data warehouses
  • Strong proficiency in T-SQL, Python, and Apache Spark
  • Practical experience with Azure DevOps, CI/CD pipelines, and infrastructure as code using Bicep, ARM, and PowerShell or Azure CLI
  • Exposure to Databricks Unity Catalog and/or Microsoft Purview
  • Experience working in Agile environments with strong communication and stakeholder engagement skills
  • Relevant Azure certifications, including AZ-900 plus one of DP-600, DP-700, or DP-203
  • Prior financial services domain experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2013

What We Do

Sabenza IT is a niche recruitment company specializing in Information Technology, SAP, Finance, and Engineering roles, with over 23 years of experience.

Similar Jobs

Mastercard Logo Mastercard

Director, Customer Success Africa and East Arabia

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Johannesburg, City of Johannesburg, Gauteng, ZAF
38800 Employees

Mastercard Logo Mastercard

Vice President, Financial Institutions Sales - Africa, Commercial Solutions

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Sandton, City of Johannesburg, Gauteng, ZAF
38800 Employees

Mastercard Logo Mastercard

Manager, Franchise Customer Onboarding and Partnership

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Johannesburg, City of Johannesburg, Gauteng, ZAF
38800 Employees

Mondelēz International Logo Mondelēz International

Digitalization & Automation lead, SSA

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
5 Locations
90000 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account