Senior Data Engineer

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in Warsaw, Warszawa, Masovian, POL
In-Office or Remote
Senior level
Software
The Role
Build and maintain large-scale data infrastructure for a real-time AdTech platform. Responsibilities include developing streaming and batch ingestion pipelines, feature tables, labeling systems, experimentation infrastructure, data-serving flows, quality frameworks, historical backfills, and privacy controls. The role requires production-grade SQL and Python engineering, Spark-based processing, cloud data warehouse expertise, point-in-time correctness, and collaboration on platform architecture and operational reliability.
Summary Generated by Built In
Company Description

Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. We are looking for a Senior Data Engineer who enjoys solving complex distributed data challenges and building production-grade ML-oriented data systems.

You will become part of a dedicated Sigma Software team developing predictive modeling and optimization capabilities for a live advertising ecosystem. The role combines large-scale event processing, streaming and batch pipelines, experimentation infrastructure, and high-throughput data engineering in a cloud-native environment.

We as a company offer the opportunity to work on impactful global products, collaborate with experienced engineers, and contribute to architecture decisions while growing your expertise in large-scale distributed systems and modern data platforms.

CUSTOMER

Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a large-scale ad exchange handling hundreds of millions of auction requests per day and is actively investing in predictive decisioning technologies to optimize advertising outcomes in real time.

PROJECT

The project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment. The platform performs real-time supply scoring and filtering, contextual performance estimation, look-alike audience generation, and multi-objective optimization under business constraints.

The solution processes massive-scale event and auction datasets and includes feature engineering pipelines, streaming and batch ingestion, experimentation infrastructure, point-in-time-correct training data generation, and ML-oriented data services with strict operational reliability and compliance requirements.

Job Description

  • Write and defend diagnostic SQL queries against large-scale production datasets
  • Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
  • Harmonize fields across independently designed datasets and maintain versioned field mappings
  • Develop point-in-time-correct feature tables and aggregation pipelines
  • Design and maintain conversion and labeling pipelines with delayed label handling
  • Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
  • Build experimentation infrastructure including traffic splitting and reporting pipelines
  • Perform large-scale historical backfills and safe reprocessing after mapping changes
  • Implement data isolation and safe-aggregation controls for advertiser data protection
  • Develop automated data quality validation frameworks
  • Collaborate closely with Customer engineers and prepare operational documentation
  • Contribute to architecture discussions and platform scalability improvements

Qualifications

  • 5+ years of experience in Data Engineering
  • At least 2 years of experience working with production ML or large-scale analytics pipelines
  • Expert-level SQL skills including window functions and incremental processing patterns
  • Strong Python skills for production-grade pipeline development
  • Hands-on experience with Spark or PySpark
  • Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
  • Experience working with cloud data warehouses at scale, preferably BigQuery
  • Strong understanding of data modeling and point-in-time correctness
  • Experience working with event-driven or clickstream datasets at very large scale
  • Experience supporting business-critical production pipelines
  • Upper-Intermediate English level or higher

WILL BE A PLUS

  • Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
  • Experience building streaming or near-real-time ingestion systems
  • Understanding of feature stores, train/serve skew, and label leakage prevention
  • Experience in AdTech or auction-based environments
  • Experience handling delayed or incomplete labels in ML systems
  • Experience with dbt or similar transformation frameworks
  • Experience delivering solutions into Customer-owned infrastructure
  • Knowledge of GDPR/CCPA-related privacy engineering practices
  • Experience with experimentation infrastructure and statistical validation pipelines
  • Experience working in hybrid cloud/on-prem Linux environments
  • Terraform and Kubernetes experience
  • Experience optimizing warehouse cost and performance

Additional Information

PERSONAL PROFILE

  • Strong analytical and problem-solving skills
  • Ownership-oriented mindset
  • Ability to work independently in a client-facing environment
  • Strong communication and documentation skills
  • Comfortable working in a fast-paced engineering environment
  • Collaborative and proactive attitude

Skills Required

  • 5+ years of experience in data engineering
  • At least 2 years of experience working with production ML or large-scale analytics pipelines
  • Expert-level SQL skills, including window functions and incremental processing patterns
  • Strong Python skills for production-grade pipeline development
  • Hands-on experience with Spark or PySpark
  • Experience designing ETL or ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
  • Experience working with cloud data warehouses at scale, preferably BigQuery
  • Strong understanding of data modeling and point-in-time correctness
  • Experience working with event-driven or clickstream datasets at very large scale
  • Experience supporting business-critical production pipelines
  • Upper-Intermediate English level or higher
  • Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
  • Experience building streaming or near-real-time ingestion systems
  • Understanding of feature stores, train/serve skew, and label leakage prevention
  • Experience in AdTech or auction-based environments
  • Experience handling delayed or incomplete labels in ML systems
  • Experience with dbt or similar transformation frameworks
  • Experience delivering solutions into customer-owned infrastructure
  • Knowledge of GDPR/CCPA-related privacy engineering practices
  • Experience with experimentation infrastructure and statistical validation pipelines
  • Experience in hybrid cloud/on-prem Linux environments
  • Terraform and Kubernetes experience
  • Experience optimizing warehouse cost and performance
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
1,516 Employees

What We Do

Sigma Software Group, an award-winning and trusted IT partner, has been serving customers for over 21 years, providing comprehensive IT solutions to various businesses, ranging from startups to established software product houses. As one of Europe's substantial IT consultancies, it brings together a dedicated workforce of over 2,100 professionals in 40 offices across 19 countries. With a diverse client base, including more than 300 enterprises, including Fortune 500 stalwarts, Sigma Software Group is a preferred choice for developing solutions that help businesses create cutting-edge products while meeting their unique needs. Sigma Software Group operates as a dynamic ecosystem of tech companies, offering 25 ready-to-implement innovative products and 40+ value-added services. Furthermore, Sigma Software Group is committed to fostering innovation through initiatives such as the Sigma Software Labs business incubator, Sigma Software University, the SID Venture Partners VC Fund, UA Tech Network, Techosystem, the European Business Association, and other collaborative efforts. Since 2015, Sigma Software Group has consistently earned recognition on the IAOP's prestigious World's Top 100 Outsourcing list. The company's accomplishments have also been acknowledged by prominent global media outlets such as Forbes, CNBC, The Times, and Reuters

Similar Jobs

Profitroom Logo Profitroom

Senior Data Engineer

Marketing Tech • Software • Hospitality
Remote
Poland
371 Employees
23K-27K Annually

Nagarro Logo Nagarro

Senior Data Engineer

Artificial Intelligence • Information Technology • Machine Learning • Software • Virtual Reality • Analytics
Remote
PL
19994 Employees

Morgan Advanced Materials, PLC Logo Morgan Advanced Materials, PLC

Senior Data Engineer

Aerospace • Energy • Industrial • Manufacturing
Remote
PL
8100 Employees

Softeta Logo Softeta

Senior Data Engineer

Information Technology • Software • Consulting
Remote
Poland
75 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account