Senior Data Engineer

Posted 4 Days Ago
Be an Early Applicant
Hiring Remotely in Kyiv, Kiev, UKR
In-Office or Remote
Senior level
AdTech
The Role
Design, build, and operate large-scale batch data pipelines and data models on a lakehouse/warehouse platform. Develop Python services/APIs, orchestrate with Airflow, ensure data correctness and performance, monitor production, and collaborate cross-functionally to deliver data products.
Summary Generated by Built In

Simulmedia is looking for an experienced and dynamic Data Engineer with a curious and creative mindset to join our Data Services team. The ideal candidate will have a strong background in Python, SQL and large-scale data pipelines. This is an opportunity to join a team of amazing engineers, data scientists, product managers and designers who are obsessed with building the most advanced TV and streaming advertising platform in the market. As a Data Engineer you will design, build and operate the data platform that powers the company: pipelines that ingest and transform very large datasets from external data partners, data models that the whole company queries, and services that make that data available to internal products. You will work on a team that empowers the other teams to use our huge amount of data efficiently. Using a large variety of technologies and tools, you will solve complicated technical problems and build solutions to make our pipelines robust and fault tolerant and our data easily accessible throughout the company.

Location: Ukraine is mandatory. Our offices are located in Kyiv and Lviv. Teams are located in Kyiv and Lviv and primarily work remotely with occasional offline meetings.

Responsibilities:

  • Design and build batch data pipelines that ingest, validate and transform multi-billion-row datasets from external data providers and internal systems
  • Model complex real-world data: dimensional models, reference data, and temporal data whose attributes change over time (e.g. slowly changing dimensions), and evolve those models safely as upstream sources change their schemas and semantics
  • Develop and operate workloads on our lakehouse platform (Databricks / Spark / Delta) and our data warehouse (Redshift), including migrating existing pipelines from the warehouse to the lakehouse
  • Orchestrate pipelines with Airflow: scheduling, dependencies, retries, backfills and alerting
  • Prove correctness, not just completion: design parity checks and reconciliation queries when replacing an existing pipeline, run large historical backfills, and investigate data discrepancies down to the row level
  • Build and maintain Python services and REST APIs that serve data to internal products
  • Optimize for performance and cost: query tuning, table design, workload management and right-sizing compute
  • Own what you ship: monitor production pipelines, participate in incident triage and root-cause analysis, and harden systems so the same failure does not happen twice
  • Collaborate cross-functionally with product managers, data scientists and stakeholders across the company to deliver on product roadmap
  • Work within an Agile team that releases cutting-edge new features regularly
  • Take a high degree of ownership and freedom to experiment with new technologies to improve our software

Qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 7+ years of work experience as a data engineer
  • Proficiency in Python and using it as the primary development language in recent years
  • Expert-level SQL: comfortable writing, reading and tuning complex analytical queries against very large tables, and debugging why two result sets disagree
  • Hands-on experience with a distributed data processing platform (Spark/Databricks strongly preferred; EMR, Snowflake or BigQuery also relevant) and with a columnar data warehouse (Redshift, Snowflake, BigQuery, ClickHouse, etc)
  • Ability to design complex data models: normalized, dimensional and temporal (slowly changing dimensions, effective-dated records, point-in-time correctness)
  • Experience with workflow orchestration tools (Airflow or similar): building DAGs, managing dependencies and running backfills
  • Experience integrating third-party data feeds: handling schema drift, late or missing deliveries, vendor data-quality defects and versioned reference data
  • Experience building REST services in Python (FastAPI, Flask, etc)
  • Experience developing, maintaining, and debugging problems in large server-side code bases
  • Working knowledge of AWS (S3, IAM, ECS or similar compute) and Docker
  • Good knowledge of engineering best practices and testing (unit test, integration test, code review, CI/CD)
  • The desire to take a high level of ownership of the things you work on
  • Ability to learn new things quickly, maintain a high bar for quality, and be pragmatic
  • Must be able to communicate with U.S based teams
  • Experience with Delta Lake / medallion lakehouse architectures is a plus
  • Experience migrating legacy pipelines between platforms with strict parity requirements is a plus
  • Experience with advertising, media or measurement industry data is a plus
  • Ability to communicate effectively with the U.S.-based teams and work 11:00 AM — 8:00 PM EEST (11:00 - 20:00).

Our Tech Stack:

  • Almost everything we run is on AWS (S3, ECS, EMR, RDS and more)
  • Python is our primary language; SQL is everywhere
  • Databricks (Spark, Delta Lake) is our lakehouse platform; Redshift and Postgres are our warehouses and operational databases
  • Airflow orchestrates our pipelines
  • Docker for packaging; GitHub Actions and Jenkins for CI/CD
  • Grafana, Sentry and OpenSearch for observability
  • Datasets measured in billions of rows
Interview Process:
  1. Pre-screening (30 mins)
  2. Technical Interview (1h)
  3. System Design (1.5h)
  4. Product Interview (30 mins)
  5. Offer

Skills Required

  • Bachelor's degree in Computer Science or equivalent experience
  • 7+ years of work experience as a data engineer
  • Proficiency in Python as primary development language
  • Expert-level SQL for complex analytical queries and debugging
  • Hands-on experience with distributed data processing platforms (Spark/Databricks, EMR, Snowflake, BigQuery)
  • Experience with columnar data warehouses (Redshift, Snowflake, BigQuery, ClickHouse)
  • Ability to design complex data models including dimensional and temporal models
  • Experience with workflow orchestration tools (Airflow or similar)
  • Experience integrating third-party data feeds and handling schema drift and data quality issues
  • Experience building REST services in Python (FastAPI, Flask, etc.)
  • Experience developing and debugging large server-side codebases
  • Working knowledge of AWS (S3, IAM, ECS, EMR, RDS) and Docker
  • Familiarity with engineering best practices and testing (unit, integration, CI/CD)
  • Ability to communicate effectively with U.S.-based teams
  • Ability to work 11:00 AM -- 8:00 PM EEST (11:00 - 20:00)
  • Experience with Delta Lake / medallion lakehouse architectures
  • Experience migrating legacy pipelines between platforms with strict parity requirements
  • Experience with advertising, media or measurement industry data
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
120 Employees
Year Founded: 2008

What We Do

Simulmedia is the leader in cross-channel TV advertising. With our TV+® platform, we deliver unparalleled reach, measurement, and results wherever audiences watch or stream. Founded In 2008, Simulmedia pioneered a data-first, digital approach to TV ad placement and optimization that changed TV advertising forever. With TV+, Simulmedia helps advertisers and agencies quickly and effectively reach viewers scattered across both linear television and CTV at guaranteed scale without wasteful duplication. Simulmedia has planned and executed successful TV campaigns for hundreds of brands, including Experian, WarnerMedia, Zelle, Disney, 1-800-FLOWERS, Monster, Electrolux, Rover, Nordstrom, King’s Hawaiian, and many more. In addition, Simulmedia allows brands to extend their reach and connect with elusive younger audiences via PlayerWON®, the first engagement and monetization platform for free-to-play PC and console video games. For more information, visit www.simulmedia.com and www.playerwon.com.

Similar Jobs

SavvyMoney Logo SavvyMoney

Senior Data Engineer

Fintech • Software • Analytics • Financial Services
Remote
27 Locations
146 Employees

Ciklum Logo Ciklum

Senior Data Engineer

Information Technology • Consulting
Remote
Ukraine
2995 Employees

Proton.ai Logo Proton.ai

Senior Data Engineer

Artificial Intelligence
Remote
38 Locations
67 Employees

N-iX Logo N-iX

Senior Data Engineer

Information Technology • Consulting
Remote
27 Locations
2135 Employees

Similar Companies Hiring

Grocery TV Thumbnail
Software • Retail • Marketing Tech • Hardware • Digital Media • AdTech
Austin, TX
56 Employees
Agentio Thumbnail
AdTech • Artificial Intelligence
New York, New York
65 Employees
ClickMint Thumbnail
AdTech • eCommerce • Marketing Tech • Generative AI
Malibu, CA
9 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account