Data Engineer (SC Cleared)

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office or Remote
Mid level
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Hands-on data engineer responsible for designing, building and maintaining scalable AWS-native data pipelines using Apache Spark/PySpark and Airflow. Duties include pipeline development, orchestration, root-cause analysis, dimensional data modelling (SCD), infrastructure as code with Terraform, containerisation with Docker, and CI/CD via GitLab. Work within strict data governance and encryption standards in a regulated government environment.
Summary Generated by Built In

Apache Spark · Python · AWS · Cloud Data Pipelines

A hands-on data engineering role within a large-scale cloud data programme, responsible for building, maintaining, and troubleshooting data pipelines using Apache Spark, PySpark, Apache Airflow, and a broad suite of AWS services. You will apply strong analytical and engineering skills to deliver trusted, well-governed data assets in a modern, cloud-native environment.

About Scrumconnect

Scrumconnect is a leading UK technology consultancy delivering digital transformation across public and private sectors, contributing to over 20% of the UK’s major citizen-facing public services. We specialise in cloud engineering, data platforms, and agile delivery, helping clients build scalable, secure, and user-centred digital solutions that create real impact.

Active SC clearance is a mandatory, non-negotiable requirement. Candidates must hold current, in-date Security Check (SC) clearance at the time of application. Sponsorship is not available. Applications without active SC clearance will not be considered.

Working arrangement:
This role is hybrid. Candidates must be willing and able to travel to the London office three days per week. Remaining days may be worked remotely from anywhere in the UK.

About the role

You will work as a Data Engineer on a complex, cloud-based data programme — designing, building, and maintaining data pipelines that process large volumes of data across a modern AWS-native stack. Using Apache Spark and PySpark for distributed data processing, Apache Airflow for orchestration, and a range of AWS services for storage, compute, and analytics, you will help deliver reliable, well-governed data assets to downstream users.

You will apply strong data analysis skills to identify root causes of data issues, work with dimensional data models and slowly changing dimensions, and implement infrastructure as code using Terraform. Familiarity with DWP engineering best practices and the ability to translate customer expectations into applied technical functionality are key to success in this role.

Key responsibilities

Data pipeline development

Build and maintain scalable data pipelines using Apache Spark and PySpark, processing and transforming large datasets across distributed cloud infrastructure.

Workflow orchestration

Configure and manage Apache Airflow DAGs for task orchestration, ensuring reliable scheduling, monitoring, and execution of data processing workflows.

Root cause analysis

Perform data analysis to identify and resolve root causes of pipeline failures and data quality issues — including reviewing EMR output logs and CloudWatch metrics.

Data modelling

Apply understanding of dimensional data models and slowly changing dimensions (SCD) to design and maintain well-structured, analytically trusted data assets.

Infrastructure as code

Provision and manage cloud infrastructure using Terraform. Containerise solutions using Docker and manage deployments through GitLab CI/CD pipelines and release tagging.

Security & encryption

Apply understanding of both server-side and client-side encryption patterns within AWS. Work within IAM policies and data governance standards appropriate to a regulated government environment.

Technical skills required

Languages & analytics
  • Python — primary language for pipeline development and data processing
  • SQL — used for querying, transformation, and validation across data stores
  • PySpark — for distributed data processing using Apache Spark on AWS EMR
  • Familiarity with basic data structures for constructing robust, scalable solutions

Data processing & orchestration

  • Apache Spark — understanding of distributed data processing architecture and execution
  • Apache Airflow — configuring DAGs and managing task orchestration at scale
  • Jupyter Notebooks — for exploratory data analysis and pipeline prototyping
  • Understanding of dimensional data models and slowly changing dimensions (SCD Types 1, 2, 3)
  • Data analysis skills to identify root cause of issues within pipelines and data assets

AWS services

  • Amazon EMR — running Spark workloads and reviewing output logs
  • Amazon Athena — ad hoc querying of data in S3
  • Amazon Textract and Comprehend — familiarity with AI/ML document extraction and NLP services
  • AWS S3, IAM, CloudWatch, EC2, ECR — core platform services used day-to-day
  • AWS console proficiency — navigating, configuring, and monitoring services
  • Understanding of server-side and client-side encryption within AWS

Infrastructure, DevOps & delivery

  • Terraform — Infrastructure as Code for provisioning and managing AWS environments
  • Docker — containerisation of data engineering solutions
  • GitLab — source code management, CI/CD pipeline configuration, release tagging, and component versioning
  • Familiarity with DWP engineering best practices
  • Ability to translate customer expectations into applied, functional technical solutions

Technology stack at a glance

PythonPySparkSQLApache SparkApache AirflowJupyter NotebooksDimensional modelling / SCDAWS EMRAmazon AthenaAWS S3AWS IAMAWS CloudWatchAWS EC2 / ECRAmazon TextractAmazon ComprehendTerraformDockerGitLab CI/CDGitLab Tags



Skills Required

  • Active Security Check (SC) clearance at time of application
  • Willingness and ability to work hybrid, attending London office three days per week
  • Python for pipeline development and data processing
  • SQL for querying, transformation, and validation
  • PySpark / Apache Spark for distributed data processing (EMR)
  • Apache Airflow for DAG configuration and orchestration
  • Experience with AWS services: EMR, Athena, S3, IAM, CloudWatch, EC2, ECR and AWS Console proficiency
  • Experience performing data analysis and root cause analysis of pipeline failures (reviewing EMR logs and CloudWatch metrics)
  • Understanding of dimensional data models and slowly changing dimensions (SCD Types 1,2,3)
  • Terraform for infrastructure as code
  • Docker for containerisation
  • GitLab for source control and CI/CD pipeline configuration, release tagging
  • Knowledge of server-side and client-side encryption patterns and working within IAM/data governance in regulated environments
  • Familiarity with Amazon Textract and Amazon Comprehend (document extraction and NLP)
  • Familiarity with DWP engineering best practices and ability to translate customer expectations into technical solutions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
300 Employees
Year Founded: 2011

What We Do

Scrumconnect Consulting is a UK national SME and award-winning software development consultancy with more than 300 consultants across the UK. It works with public-sector clients to build impactful digital services that improve millions of lives, combining design, engineering, and AI expertise. The company emphasizes experienced consultants, collaboration, knowledge-sharing, problem-solving, and a culture focused on sustaining innovation for its customers.

Similar Jobs

Scrumconnect Limited Logo Scrumconnect Limited

Test Engineer

Artificial Intelligence • Information Technology • Software • Consulting
In-Office or Remote
2 Locations
300 Employees

HiBob Logo HiBob

Team Lead

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
United Kingdom
1350 Employees

Atlassian Logo Atlassian

Head of Sales Development, Strategic & Enterprise

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
London, Greater London, England, GBR
11000 Employees

Block Logo Block

Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
London, Greater London, England, GBR
12000 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account