We are looking for Data Engineer in Kraków, Poland who will own the company data lake — the system of record for millions of financial events a day (bets, wallet movements, live odds) across 12M+ active users. You decide how that data lands, is stored, retained, and governed on S3 + Snowflake/Databricks, so analytics, finance, and regulators all see accurate, reconciled data with zero drift from source.
Domain: Regulated iGaming / wallet & ledger data. Audit-heavy: regulators, finance and analytics all consume the same tables. Millions of financial events per day, terabyte-plus scale.
What you'll be doing:
- Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.
- Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.
- Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
- Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.
- Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.
- Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.
- Monitor ingestion health, data anomalies, and cloud storage/compute spend.
Must-have:
- 5+ years in data engineering, with real ownership of a large-scale data lake or lakehouse.
- Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.
- Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
- Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.
- CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.
- Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from. Critical for financial ledgers.
- Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.
- Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.
- Cloud object storage — S3 or GCS, plus storage tiering and archival.
- Python and an orchestrator — Airflow or Dagster, as tools.
Location & work model:
Kraków, Poland. Hybrid — 2 days per week from the office.
Skills Required
- 5+ years of data engineering experience with ownership of a large-scale data lake or lakehouse
- Experience with bronze, silver, and gold lakehouse architecture, open table formats such as Iceberg, Delta, or Hudi, and schema evolution
- Experience optimizing data layout and queries at terabyte-plus scale, including partitioning, compaction, file sizing, and Trino, Athena, or Snowflake performance
- Production experience with Snowflake, Databricks, or BigQuery
- Experience with CDC and streaming ingestion using Kafka, Debezium, or equivalent technologies
- Strong SQL and data modeling skills, including relational and OLTP systems
- Experience implementing freshness SLAs, drift detection, and reconciliation against source wallet or ledger systems
- Experience with data governance, catalogs, lineage, row and column access control, PII masking, retention, and GDPR deletion
- Experience with cloud object storage such as S3 or GCS, including storage tiering and archival
- Experience with Python and a workflow orchestrator such as Airflow or Dagster
What We Do
Commit is a global tech services company with offices in Israel, US, Canada, UK, and Europe. The company was founded in 2005 and has over 700 multi-disciplinary innovation experts who serve a broad range of companies, from small startups to large enterprises in multiple business sectors. Commit specializes in advanced technologies and applications with dedicated practices in Cloud, GenAI, Software, IoT, Big Data, Cyber, Collaboration, Data center migration projects, and more. Commit offers innovative, end-to-end technology solutions by developing custom software and IoT platforms for clients looking to build their next-gen products within the modern ICT world. Commit’s complete and comprehensive engineering powerhouse of resources, and proprietary Flexible R&D methodology helps transform its clients’ technology visions into high-quality products while reducing costs and improving time-to-market.









