Data Engineer

Posted One Month Ago
Be an Early Applicant
Mbarara, Kamukuzi Division, Mbarara Municipality, Mbarara, UGA
In-Office
Junior
Social Impact • Agriculture
The Role
Build and maintain batch and streaming data pipelines in Databricks using Python, PySpark, SQL, Delta Lake, and Medallion Architecture. Develop warehouse layers, data models, observability, validation, monitoring, and alerting systems. Support ML feature engineering and training-data workflows by integrating field data and model outputs. Collaborate with engineers, data scientists, field teams, and program staff while documenting architectures, transformations, data dictionaries, and runbooks.
Summary Generated by Built In

Job Title: Data Engineer

Department/Group: Venn

Reporting To: Senior Data Scientist

Years of Experience: 1-3 years

Location: Mbarara

Travel Required: up to 20%

About Raising The Village

Raising The Village's mission is to build and strengthen pathways out of ultra-poverty through an approach guided by data, designed with communities, and scaled by partnerships. We work with last-mile communities to implement solutions that raise incomes, sustain impact, and chart sustainable pathways to possibility within 24 months. To date, we have supported more than two million people in Uganda, Rwanda, Tanzania and the Democratic Republic of Congo, with support from our partners and our team in North America. Find out more about our programs and impact at www.raisingthevillage.org.

The VENN department is the data and technology backbone of our organization, connecting advanced analytics and custom software tools with field implementation to ensure data-informed decision-making at every level.

The Opportunity

RTV's data infrastructure is in an active build phase, with a roadmap that is being deliberately accelerated to keep pace with the organization's expanding program footprint across Uganda, Rwanda, and the Democratic Republic of Congo. The data engineering team is growing, and this hire is part of that growth.

You will join a small, senior-led team at a moment when foundational architecture decisions are still being made, pipelines are being built from the ground up, and the systems you contribute to will directly shape how RTV measures impact, deploys AI tools, and scales programmatic reach to millions of people living in ultra-poverty.

Role Description

The Data Engineer is a core builder on RTV's expanding data infrastructure team, reporting directly to the Senior Data Scientist within the VENN department, and responsible for designing, developing, and maintaining the pipelines, warehouse layers, and data quality systems that power RTV's programmatic analytics, machine learning platforms, and field evaluation tools. Operating at the intersection of data platform engineering, ML infrastructure support, and field data integration, this role works across a fast-moving roadmap that spans batch and streaming ingestion, ELT pipeline development, Delta Lake architecture, observability frameworks, and the integration of structured field data with AI model outputs.

Key Responsibilities

Pipeline Development & Delivery

  • Design and deliver batch and streaming data pipelines that ingest data from field collection platforms (SurveyCTO, ArcGIS, custom mobile apps) into RTV's Databricks warehouse, working at pace against an accelerated infrastructure roadmap.
  • Implement ELT and ETL workflows in PySpark and Databricks SQL, applying Medallion Architecture (Bronze, Silver, Gold) principles to produce clean, versioned, and consumption-ready data layers.
  • Contribute actively to sprint-based delivery cycles, taking ownership of pipeline workstreams end-to-end from design through deployment and monitoring.

Delta Lake & Warehouse Architecture

  • Build and maintain Delta Lake table structures with strong schema enforcement, ACID-compliant write patterns, and time travel capabilities to support auditability across program datasets.
  • Contribute to data modelling decisions including star schema design, SCD patterns, and denormalization trade-offs.
  • Support the evolution of Unity Catalog governance structures including lineage tracking, access controls, and dataset documentation as the warehouse scales across new program domains and geographies.

Data Observability & Quality

  • Support the maintenance of data observability frameworks across all pipeline stages, covering data freshness, volume anomalies, schema drift detection, and SLA monitoring.
  • Build validation and quality checks at ingestion and transformation layers using tools such as Great Expectations, dbt tests, or Databricks-native monitoring capabilities.
  • Contribute to structured logging, alerting, and incident response practices that give the team fast, reliable visibility into pipeline health across all environments.

ML & AI Pipeline Support

  • Integrate structured household data, image classification outputs, and ML model predictions into unified warehouse layers for consumption by Data Scientists and the WorkMate AI platform.
  • Collaborate with ML Engineers and Data Scientists to build and maintain feature engineering pipelines and training data preparation workflows that support RTV's computer vision and adoption scoring systems.

Collaboration & Documentation

  • Work closely with the technical team on roadmap prioritization, architectural decisions, and engineering standards as the team scales.
  • Partner with Software Engineers, Data Scientists, field evaluation teams, and program staff to understand data requirements and translate them into reliable, well-documented pipeline solutions.
  • Maintain thorough documentation of pipeline architectures, transformation logic, data dictionaries, and runbooks to support team growth and organizational knowledge continuity.

Technical Requirements

Education & Experience: A Bachelor's degree in Computer Science, Software Engineering, Data Engineering, Information Systems, Statistics, or a related quantitative field is preferred. Equivalent practical experience through demonstrable project work, open-source contributions, or bootcamp training is equally welcome. Clear evidence of building and shipping production-grade pipelines, including demonstrable ownership from design through deployment with specific examples of integrating, moving, and transforming data at a meaningful scale.

Technical Skills: Candidates should have strong hands-on experience with Python (including Pandas), SQL, and building or supporting ETL/ELT pipelines. Experience with PySpark, Databricks, or similar modern data platforms is highly valued. Exposure to Delta Lake, Unity Catalog, Medallion Architecture, streaming pipelines, data observability, AWS, and BI tools is considered an asset. The successful candidate will continue to build depth in these areas while working closely with the Senior Data Scientist and broader technical team.

We encourage women, people with disabilities and minority groups to apply for this position. RTV is committed to equal opportunities and diversity of perspective at the workplace. 

Disclaimer: Raising the Village DOES NOT charge any kind of FEE(s) at whichever stage of the recruitment process

Skills Required

  • Demonstrable experience building and shipping production-grade data pipelines, including ownership from design through deployment and integrating, moving, and transforming data at meaningful scale
  • Proficiency in Python, including Pandas
  • Proficiency in PySpark
  • Proficiency in advanced SQL for data transformation at scale
  • Experience working in a Databricks-first environment with Delta Lake, Unity Catalog, and Medallion Architecture
  • Experience designing ELT and ETL workflows
  • Experience building batch and Spark Structured Streaming pipelines
  • Experience with data observability, validation, monitoring, and alerting frameworks
  • Working knowledge of Git-based collaborative workflows
  • Working knowledge of AWS core services
  • Ability to produce visualization-ready datasets for Power BI, Tableau, or Python visualization libraries
  • Bachelor's degree in Computer Science, Software Engineering, Data Engineering, Information Systems, Statistics, or a related quantitative field
  • Equivalent practical experience through project work, open-source contributions, or bootcamp training
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
366 Employees
Year Founded: 2012

What We Do

Raising The Village is a Canadian not-for-profit organization dedicated to ending ultra-poverty in Sub-Saharan Africa. The organization partners with last-mile communities to increase household income and production through a data-driven, holistic model. By increasing agricultural yields, diversifying income, and alleviating barriers to economic participation, they build sustainable pathways out of poverty across Uganda, Rwanda, and the Democratic Republic of Congo.

Similar Jobs

In-Office or Remote
2 Locations
25000 Employees
In-Office or Remote
2 Locations
25000 Employees

Airtel Africa plc Logo Airtel Africa plc

Lead - Digital Channels

Information Technology • Mobile
In-Office or Remote
2 Locations
16118 Employees
In-Office
Mbarara, Kamukuzi Division, Mbarara Municipality, Mbarara, UGA
366 Employees

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Playground (tryplayground.com) Thumbnail
Kids + Family • Payments • Social Impact • Software
New York City, New York
80 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account