Mid Data Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Poland
Remote or Hybrid
Mid level
Information Technology • Software
The Role
Own and evolve a production lakehouse supporting millions of daily financial events. Design bronze, silver, and gold layers; build CDC streaming pipelines; optimize storage and queries; manage retention, archival, reconciliation, and data correctness. Establish governance through cataloging, lineage, access controls, masking, encryption, and audit trails. Monitor ingestion health, anomalies, and cloud costs while supporting reliable, traceable data for Analytics, Finance, and Regulatory teams.
Summary Generated by Built In
Description

We’re looking for a Middle Data Engineer to take ownership of our data lake — the system of record for millions of financial events every day, including bets, wallet transactions, and live odds, serving 12M+ active users.You’ll be responsible for designing and maintaining reliable data pipelines and defining how data is ingested, stored, retained, reconciled, and governed across AWS S3 and Snowflake/Databricks. Your work will ensure that Analytics, Finance, and Regulatory teams have access to accurate, consistent, and fully traceable data — with zero drift from source systems.

Location: Kraków, Poland. Hybrid — 2 days per week from the office.

What you will do:

  • Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.
  • Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.
  • Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
  • Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.
  • Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.
  • Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.
  • Monitor ingestion health, data anomalies, and cloud storage/compute spend.
Requirements

Must-have:

  • 3+ years hands-on in a production lakehouse environment.
  • Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.
  • Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
  • Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.
  • CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.
  • Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from. Critical for financial ledgers.
  • Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.
  • Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.
  • Cloud object storage — S3 or GCS, plus storage tiering and archival.
  • Python and an orchestrator — Airflow or Dagster, as tools.

Nice to have:

  • Fintech, iGaming, or another regulated, audit-heavy environment.
  • Cost monitoring / FinOps for storage and compute spend.
  • Hudi specifically; Dagster specifically.
  • Immutability / WORM regulatory retention.

Skills Required

  • 3+ years of hands-on experience in a production lakehouse environment
  • Experience with lakehouse architecture, including bronze/silver/gold layers, open table formats, and schema evolution
  • Experience with data layout and query optimization at TB+ scale, including partitioning, compaction, file sizing, and query performance
  • Production experience with Snowflake, Databricks, or BigQuery
  • CDC and streaming ingestion experience with Kafka, Debezium, or equivalent
  • Strong SQL and data modeling skills
  • Experience with freshness SLAs, drift detection, and reconciliation against source systems
  • Experience with data catalogs, lineage, row/column access control, PII masking, retention, and GDPR deletion
  • Experience with cloud object storage such as S3 or GCS, including storage tiering and archival
  • Python and experience with an orchestrator such as Airflow or Dagster
  • Experience in fintech, iGaming, or another regulated, audit-heavy environment
  • Cost monitoring or FinOps experience for storage and compute spend
  • Specific experience with Hudi or Dagster
  • Experience with immutability or WORM regulatory retention
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Petaling Jaya
399 Employees
Year Founded: 2005

What We Do

Commit is a global tech services company with offices in Israel, US, Canada, UK, and Europe. The company was founded in 2005 and has over 700 multi-disciplinary innovation experts who serve a broad range of companies, from small startups to large enterprises in multiple business sectors. Commit specializes in advanced technologies and applications with dedicated practices in Cloud, GenAI, Software, IoT, Big Data, Cyber, Collaboration, Data center migration projects, and more. Commit offers innovative, end-to-end technology solutions by developing custom software and IoT platforms for clients looking to build their next-gen products within the modern ICT world. Commit’s complete and comprehensive engineering powerhouse of resources, and proprietary Flexible R&D methodology helps transform its clients’ technology visions into high-quality products while reducing costs and improving time-to-market.

Similar Jobs

Xebia Logo Xebia

Senior Data Engineer

Artificial Intelligence • Cloud • Information Technology • Software • Consulting • Data Privacy
Remote
3 Locations
3254 Employees

Pfizer Logo Pfizer

Staff Software Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees

GitLab Logo GitLab

Architect

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
5 Locations
2500 Employees
168K-238K Annually

Capco Logo Capco

Devsecops Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account