We are looking for a Senior Data Engineer to join a high-performing data engineering team at a leading global company operating at the intersection of sports, technology, and digital commerce.
In this role, you will take ownership of the systems responsible for ingesting, transforming, validating, and publishing data across a large-scale data ecosystem. You will work at the data ingestion boundary, where information from multiple internal and external sources enters the platform, ensuring it is accurate, reliable, and ready to power downstream products, analytics, and services.
You will be responsible for building and maintaining scalable data pipelines, solving complex identity and entity-matching challenges, identifying and addressing data-quality issues, and improving the reliability and observability of data flows across the organization.
This is a highly hands-on engineering role that combines AWS data engineering, Python, SQL, event-driven architectures, data quality, and identity resolution. You will work closely with engineering, product, analytics, and downstream platform teams to troubleshoot complex data challenges, improve existing systems, and build reliable solutions that support data-driven products at scale.
What You'll DoOwn and evolve data ingestion pipelines that bring data from multiple external and internal sources into the platform, including ingestion, cleaning, curation, entity resolution, and event-driven publishing.
Design, build, and maintain scalable batch and event-driven data workflows using Python, SQL, and AWS.
Solve complex identity resolution, entity matching, and deduplication challenges across multiple data sources, ensuring reliable and consistent entity mappings.
Build and maintain data quality and validation frameworks, including freshness monitoring, null-rate and consistency checks, and schema-drift detection.
Monitor data at the ingestion boundary and proactively identify, troubleshoot, and resolve issues before they impact downstream products and services.
Partner with engineering teams responsible for event infrastructure and downstream identity services to trace data and events end-to-end.
Collaborate with Product, Assessments, Analytics, and Engineering teams to understand how data is consumed downstream and ensure pipelines meet evolving product and business requirements.
Work with AWS services such as Glue, Athena, S3, and DynamoDB to build and operate reliable, scalable data infrastructure.
Automate infrastructure and pipeline changes using Infrastructure as Code, primarily Terraform or AWS CDK.
Participate in production incident response, quickly assessing impact, identifying root causes, and implementing effective remediation.
Continuously improve the reliability, observability, scalability, and maintainability of data pipelines and platform infrastructure.
5+ years of experience building and operating production-grade data pipelines and data infrastructure.
Strong experience with both batch processing and event-driven architectures.
Hands-on experience with AWS Glue, Athena, and S3-based data lakes, including layered or medallion-style data transformations.
Strong proficiency in Python and SQL for data processing, transformation, and analysis.
Experience solving identity resolution, entity matching, and deduplication problems, including exact and fuzzy matching approaches.
Understanding of durable identifiers and identity-mapping strategies, including the challenges associated with maintaining consistent first-seen or locked mappings over time.
Experience working with event schemas and schema-registry-backed contracts, such as Protobuf, and an understanding of the impact of schema evolution on downstream consumers.
Strong troubleshooting and production incident-response skills, with the ability to assess impact, identify root causes, and drive issues through resolution.
Ability to understand and debug code written in a functional or concurrent programming language, such as Elixir.
Strong communication and collaboration skills, with the ability to work effectively across engineering, product, analytics, and other technical teams.
Experience with Elixir/Phoenix or other BEAM-based concurrent processing frameworks.
Experience with DynamoDB-backed identity, lookup, or matching services.
Experience with the Snowflake ecosystem, including data modeling, Snowpipe, Streams, and Tasks.
Experience working with sports data providers, licensed data providers, or other complex external data ecosystems.
Experience implementing data observability, including freshness and staleness alerts, null-rate monitoring, schema-drift detection, and data-quality dashboards.
Experience managing infrastructure through Terraform or AWS CDK.
Languages: Python, SQL, familiarity with Elixir
AWS: Glue, Athena, S3, DynamoDB
Data & Streaming: Data Lakes, Event-Driven Architecture, Protobuf, Schema Registries
Infrastructure: Terraform, AWS CDK
Data Platforms: Snowflake
Engineering Practices: Data Quality, Data Observability, Identity Resolution, Entity Matching, Incident Response
Skills Required
- 5+ years of experience building and operating production-grade data pipelines
- Experience with batch data processing and event-driven architectures
- Hands-on experience with AWS Glue, Athena, and S3-based data lakes
- Strong proficiency in Python and SQL
- Experience with identity resolution, entity matching, or deduplication
- Understanding of durable identifiers and first-seen or locked identity mappings
- Experience with event schemas and schema-registry-backed contracts such as Protobuf
- Strong production troubleshooting and incident-response skills
- Ability to understand and debug functional or concurrent programming languages such as Elixir
- Strong communication and cross-functional collaboration skills
- Experience with Elixir/Phoenix or another BEAM-based concurrent processing framework
- Experience with DynamoDB-backed identity, lookup, or matching services
- Experience with Snowflake, including data modeling, Snowpipe, Streams, and Tasks
- Experience with sports data providers or licensed-content data ecosystems
- Experience implementing data observability and data-quality dashboards
- Experience with Terraform or AWS CDK
What We Do
Where startups soar! We're not just building startups - we're crafting the future. From ideation to execution, we provide the blueprint for startup success. We specialize in nurturing startups from the ground up, and supercharging existing ones with our top-tier technical talent solutions. Join us at Ryz Labs, where we turn promising ideas into thriving businesses. Let's shape the future of innovation together!








