The Role
Own the architecture, reliability, governance, and evolution of the Databricks lakehouse. Design batch and streaming pipelines, develop SQL/Python/PySpark transformations, productionize analytical and AI prototypes, and implement testing, orchestration, monitoring, data-quality controls, governance, and access management. Collaborate with analysts, Engineering, and Product to deliver scalable, cost-effective data solutions for clinical trial insights.
Summary Generated by Built In
At CluePoints, we're redefining how clinical trials are run. As a leading providerof Risk-Based Quality Management and Data Quality Oversight software, we useadvanced statistics, artificial intelligence and machine learning to improve thequality, accuracy and integrity of clinical trial data — turning artificialintelligence into human intelligence.We're an ambitious, fast-growing technology company with a dynamic anddiverse international team. Collaboration, flexibility and continuous learning arepart of our DNA. Guided by our values of Care, Passion and Smart Disruption,we're united by a shared mission: to create smarter ways to run efficient clinicaltrials and deliver AI-powered insights that improve human outcomes worldwide.
The Role
We're looking for a hands-on Senior Data Engineer to own the technical design, reliability and evolution of our Databricks data foundation. You'll work alongside Clinical Data Insights Analysts — colleagues with strong business, clinical and BI expertise who also write SQL and Python — complementing their skills with deeper engineering, automation and production-grade rigour.
Job requirements
Nice to Have:
Tech Stack:
Databricks · Delta Lake · Unity Catalog · Lakeflow Jobs & Declarative Pipelines · Medallion architecture · SQL / Python / PySpark · Microsoft Azure · Git / CI/CD · Databricks Asset Bundles · Databricks AI/BI · Genie
Job responsibilities
Benefits
🇧🇪 What We Offer – Belgium
Equal Opportunities & GDPR Notice
CluePoints is an equal opportunities employer. We value and respect diversity in our workforce and do not tolerate discrimination based on gender, age, disability, ethnic origin, religion, sexual orientation, or any other protected ground under Belgian law.
Personal data collected as part of your application will be processed in compliance with the EU GDPR and Belgian data protection legislation.
You have the right to access, correct, or delete your personal data at any time by contacting [email protected].
The Role
We're looking for a hands-on Senior Data Engineer to own the technical design, reliability and evolution of our Databricks data foundation. You'll work alongside Clinical Data Insights Analysts — colleagues with strong business, clinical and BI expertise who also write SQL and Python — complementing their skills with deeper engineering, automation and production-grade rigour.
Job requirements
- 5+ years in data engineering or data-platform engineering.
- Strong hands-on Databricks experience (or comparable cloud lakehouse platform).
- Advanced SQL and strong Python / PySpark skills.
- Production-grade pipeline experience — design, build, test, deploy and maintain.
- Solid understanding of Delta Lake, medallion architecture, and OLAP/OLTP transformation.
- Experience with data governance, Unity Catalog (or equivalent), automated quality controls and CI/CD.
- At least one major cloud platform: Azure, AWS or GCP.
- Practical use of generative AI in engineering workflows, with solid critical judgement on its outputs.
- High autonomy — you take ownership from design through deployment, monitoring and maintenance.
- Collaborative by nature; comfortable working directly alongside analysts who also write code.
Nice to Have:
- Lakeflow Jobs / Lakeflow Declarative Pipelines; Databricks Asset Bundles.
- Databricks AI/BI, Genie, conversational analytics or AI agent experience.
- Databricks Data Engineer certification.
- Experience in a regulated environment (clinical trials, life sciences or similar).
Tech Stack:
Databricks · Delta Lake · Unity Catalog · Lakeflow Jobs & Declarative Pipelines · Medallion architecture · SQL / Python / PySpark · Microsoft Azure · Git / CI/CD · Databricks Asset Bundles · Databricks AI/BI · Genie
Job responsibilities
- Own the architecture, reliability and governance of the team's Databricks lakehouse (Delta Lake, medallion architecture, Unity Catalog).
- Design and operate batch and streaming ingestion pipelines from the CluePoints platform and other business systems (e.g. Zendesk).
- Co-develop SQL, Python and PySpark transformations with the analyst team; review and optimize for performance, scalability and cost.
- Productionize analytical and AI prototypes — adding testing, orchestration, monitoring, error handling and deployment controls.
- Build automated data-quality controls, monitoring and alerting; investigate and resolve pipeline incidents.
- Implement technical governance: naming conventions, metadata, lineage, tagging and access controls via Unity Catalog.
- Apply AI and automation across the data lifecycle — code development, documentation, metadata classification and inconsistency detection.
- Collaborate with Engineering on source-system access and with Product when Databricks insights are candidates for platform integration.
Benefits
🇧🇪 What We Offer – Belgium
- Health Insurance through Alan (100% hospitalisation cover, 80% ambulatory and dental)
- Mobility Budget for eco-transport, housing, or car allowance (flexible 3-pillar system)
- Group Insurance Plan with 6–12% employer pension contribution based on seniority
- Meal Vouchers (€8/day) and Eco Vouchers for sustainable purchases
- A hub-based hybrid model that blends flexibility with purpose — connecting teams through collaboration, learning, and a vibrant social culture.
Equal Opportunities & GDPR Notice
CluePoints is an equal opportunities employer. We value and respect diversity in our workforce and do not tolerate discrimination based on gender, age, disability, ethnic origin, religion, sexual orientation, or any other protected ground under Belgian law.
Personal data collected as part of your application will be processed in compliance with the EU GDPR and Belgian data protection legislation.
You have the right to access, correct, or delete your personal data at any time by contacting [email protected].
Skills Required
- 5+ years of experience in data engineering or data-platform engineering
- Strong hands-on Databricks experience or comparable cloud lakehouse platform experience
- Advanced SQL skills
- Strong Python and PySpark skills
- Production-grade pipeline experience, including design, development, testing, deployment, and maintenance
- Understanding of Delta Lake, medallion architecture, and OLAP/OLTP transformation
- Experience with data governance, Unity Catalog or equivalent, automated quality controls, and CI/CD
- Experience with at least one major cloud platform: Azure, AWS, or GCP
- Practical use of generative AI in engineering workflows and ability to critically evaluate its outputs
- Ability to work autonomously from design through deployment, monitoring, and maintenance
- Collaborative approach and comfort working with analysts who write code
- Experience with Lakeflow Jobs, Lakeflow Declarative Pipelines, or Databricks Asset Bundles
- Experience with Databricks AI/BI, Genie, conversational analytics, or AI agents
- Databricks Data Engineer certification
- Experience in a regulated environment such as clinical trials or life sciences
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
CluePoints provides pharmaceutical companies and contract research organizations (CROs) with technology for identifying, visualizing, managing, and documenting risks that could affect clinical-trial outcomes. Its mission is to give clinical research organizations a better way to oversee risk, helping them recognize, assess, and document potential problems that might undermine trial success and strengthen the quality and reliability of clinical-trial decision-making.







