This is a remote position.
We are looking for an experienced Data Engineer to support the migration and integration of data from multiple enterprise and laboratory systems. The project is primarily focused on dbt, advanced SQL, and data modelling rather than Databricks or large-scale PySpark processing.
The successful candidate will design new data models, develop complex transformations, and consolidate data from several source databases into a unified AWS-based data platform.
- Design logical and physical data models based on technical and business requirements.
- Develop, test, and maintain complex data transformations using dbt and SQL.
- Migrate and consolidate data from multiple databases and source systems.
- Use Trino/Presto to query, join, and transform data distributed across different sources.
- Integrate data from enterprise and laboratory systems such as SAP, GLIMS, and Veeva.
- Build reusable, maintainable, and well-documented transformation models.
- Validate migrated data and investigate data-quality or consistency issues.
- Optimize complex SQL queries and transformation processes.
- Support data lineage, traceability, integrity, and documentation.
- Work with data stored or processed within AWS, particularly Amazon S3 and Athena.
- Participate in Agile delivery, code reviews, testing, and CI/CD activities.
- Collaborate with data architects, analysts, engineers, and pharmaceutical business stakeholders.
- Develop new solutions rather than only maintaining existing pipelines.
Requirements
- Strong hands-on experience with dbt.
- Advanced proficiency in SQL, including:
- Complex joins and transformations
- CTEs and window functions
- Query optimization
- Data reconciliation and validation
- Complex joins and transformations
- Practical experience with data modelling, including the ability to design a model from business or technical requirements.
- Experience migrating and consolidating data from multiple databases.
- Experience with distributed query engines such as:
- Trino
- Presto
- Trino
- Working knowledge of AWS data services, particularly:
- Amazon S3
- Amazon Athena
- AWS Glue
- Amazon S3
- Experience building reliable, production-ready data transformation pipelines.
- Understanding of relational databases and data warehousing concepts.
- Python, Scala, or PySpark for scripting, automation, or supplementary data transformation.
- Experience with PostgreSQL or another relational target database.
- Familiarity with Git and CI/CD deployment pipelines.
- Experience working in an Agile delivery environment.
- Knowledge of data quality, lineage, governance, and integrity principles.
Experience integrating data from any of the following is valuable:
- SAP ERP
- GLIMS or other LIMS platforms
- Veeva
- Pharmaceutical manufacturing or laboratory systems
Experience with GxP regulations and pharmaceutical data-integrity requirements is preferred but not necessarily essential.
Candidates should meet one of the following:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, with at least 4 years of relevant technical experience; or
- At least 8 years of equivalent experience in data engineering, data modelling, data integration, or cloud analytics without a degree.
The ideal candidate is a SQL-focused Data Engineer with strong practical experience in dbt and data modelling. They should be comfortable combining data from several databases, designing new data structures, and implementing complex transformations with Trino/Presto in an AWS environment.
Benefits
- Location: European Union
- Contract Type: Freelance / Contract
- Start date: Summer, 2026
- Time Allocation: 40 hours/week
- Global Pharmaceutical Company in Prague
Skills Required
- Strong hands-on experience with dbt
- Advanced SQL proficiency, including complex joins, transformations, CTEs, window functions, query optimization, reconciliation, and validation
- Practical data modeling experience, including designing models from business or technical requirements
- Experience migrating and consolidating data from multiple databases
- Experience with distributed query engines such as Trino or Presto
- Working knowledge of Amazon S3, Amazon Athena, and AWS Glue
- Experience building reliable, production-ready data transformation pipelines
- Understanding of relational databases and data warehousing concepts
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, with at least 4 years of relevant technical experience
- At least 8 years of equivalent experience in data engineering, data modeling, data integration, or cloud analytics without a degree
- Python, Scala, or PySpark for scripting, automation, or supplementary data transformation
- Experience with PostgreSQL or another relational target database
- Familiarity with Git and CI/CD deployment pipelines
- Experience working in an Agile delivery environment
- Knowledge of data quality, lineage, governance, and integrity principles
- Experience integrating SAP ERP, GLIMS or other LIMS platforms, Veeva, or pharmaceutical manufacturing and laboratory systems
- Experience with GxP regulations and pharmaceutical data-integrity requirements
What We Do
futureproof s.r.o. is a Czech consulting and staffing firm focused on data, analytics, cybersecurity, and IT infrastructure. It provides contract and permanent staffing, team augmentation, time-and-material resources, and specialist or lead placements, while also offering expert consulting through a network of architects and project leaders. The company emphasizes niche technical expertise, trusted relationships, continuous learning, and long-term value for clients.
%20copy.jpg)








