Data Engineer (AWS, DBT, Trino/Databricks)

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in CZE
Remote
Mid level
Information Technology • Analytics • Consulting • Cybersecurity
The Role
Design and implement data models and complex dbt/SQL transformations for migrating and consolidating enterprise and laboratory data into an AWS platform. Query distributed sources with Trino/Presto, integrate systems such as SAP, GLIMS, and Veeva, validate data quality, optimize pipelines, and maintain lineage and documentation. Collaborate with architects, analysts, engineers, and pharmaceutical stakeholders in an Agile environment, supporting testing, code reviews, and CI/CD.
Summary Generated by Built In

This is a remote position.

We are looking for an experienced Data Engineer to support the migration and integration of data from multiple enterprise and laboratory systems. The project is primarily focused on dbt, advanced SQL, and data modelling rather than Databricks or large-scale PySpark processing.

The successful candidate will design new data models, develop complex transformations, and consolidate data from several source databases into a unified AWS-based data platform.

Key Responsibilities
  • Design logical and physical data models based on technical and business requirements.
  • Develop, test, and maintain complex data transformations using dbt and SQL.
  • Migrate and consolidate data from multiple databases and source systems.
  • Use Trino/Presto to query, join, and transform data distributed across different sources.
  • Integrate data from enterprise and laboratory systems such as SAP, GLIMS, and Veeva.
  • Build reusable, maintainable, and well-documented transformation models.
  • Validate migrated data and investigate data-quality or consistency issues.
  • Optimize complex SQL queries and transformation processes.
  • Support data lineage, traceability, integrity, and documentation.
  • Work with data stored or processed within AWS, particularly Amazon S3 and Athena.
  • Participate in Agile delivery, code reviews, testing, and CI/CD activities.
  • Collaborate with data architects, analysts, engineers, and pharmaceutical business stakeholders.
  • Develop new solutions rather than only maintaining existing pipelines.


Requirements
  • Strong hands-on experience with dbt.
  • Advanced proficiency in SQL, including:
    • Complex joins and transformations
    • CTEs and window functions
    • Query optimization
    • Data reconciliation and validation
  • Practical experience with data modelling, including the ability to design a model from business or technical requirements.
  • Experience migrating and consolidating data from multiple databases.
  • Experience with distributed query engines such as:
    • Trino
    • Presto
  • Working knowledge of AWS data services, particularly:
    • Amazon S3
    • Amazon Athena
    • AWS Glue
  • Experience building reliable, production-ready data transformation pipelines.
  • Understanding of relational databases and data warehousing concepts.
Additional Relevant Skills
  • Python, Scala, or PySpark for scripting, automation, or supplementary data transformation.
  • Experience with PostgreSQL or another relational target database.
  • Familiarity with Git and CI/CD deployment pipelines.
  • Experience working in an Agile delivery environment.
  • Knowledge of data quality, lineage, governance, and integrity principles.
Domain Experience

Experience integrating data from any of the following is valuable:

  • SAP ERP
  • GLIMS or other LIMS platforms
  • Veeva
  • Pharmaceutical manufacturing or laboratory systems

Experience with GxP regulations and pharmaceutical data-integrity requirements is preferred but not necessarily essential.

Qualifications

Candidates should meet one of the following:

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, with at least 4 years of relevant technical experience; or
  • At least 8 years of equivalent experience in data engineering, data modelling, data integration, or cloud analytics without a degree.
Ideal Candidate

The ideal candidate is a SQL-focused Data Engineer with strong practical experience in dbt and data modelling. They should be comfortable combining data from several databases, designing new data structures, and implementing complex transformations with Trino/Presto in an AWS environment.



Benefits
  • Location: European Union
  • Contract Type: Freelance / Contract
  • Start date: Summer, 2026
  • Time Allocation: 40 hours/week
  • Global Pharmaceutical Company in Prague​


Skills Required

  • Strong hands-on experience with dbt
  • Advanced SQL proficiency, including complex joins, transformations, CTEs, window functions, query optimization, reconciliation, and validation
  • Practical data modeling experience, including designing models from business or technical requirements
  • Experience migrating and consolidating data from multiple databases
  • Experience with distributed query engines such as Trino or Presto
  • Working knowledge of Amazon S3, Amazon Athena, and AWS Glue
  • Experience building reliable, production-ready data transformation pipelines
  • Understanding of relational databases and data warehousing concepts
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, with at least 4 years of relevant technical experience
  • At least 8 years of equivalent experience in data engineering, data modeling, data integration, or cloud analytics without a degree
  • Python, Scala, or PySpark for scripting, automation, or supplementary data transformation
  • Experience with PostgreSQL or another relational target database
  • Familiarity with Git and CI/CD deployment pipelines
  • Experience working in an Agile delivery environment
  • Knowledge of data quality, lineage, governance, and integrity principles
  • Experience integrating SAP ERP, GLIMS or other LIMS platforms, Veeva, or pharmaceutical manufacturing and laboratory systems
  • Experience with GxP regulations and pharmaceutical data-integrity requirements
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2015

What We Do

futureproof s.r.o. is a Czech consulting and staffing firm focused on data, analytics, cybersecurity, and IT infrastructure. It provides contract and permanent staffing, team augmentation, time-and-material resources, and specialist or lead placements, while also offering expert consulting through a network of architects and project leaders. The company emphasizes niche technical expertise, trusted relationships, continuous learning, and long-term value for clients.

Similar Jobs

Remote
Czech Republic
575 Employees
948K-1M Annually

Pfizer Logo Pfizer

Director R&D EHS Program Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees
177K-294K Annually

SailPoint Logo SailPoint

Consultant

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
2461 Employees

Mondelēz International Logo Mondelēz International

Program Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
9 Locations
90000 Employees
4K-4K Annually

Similar Companies Hiring

Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account