Data Engineer

Posted Yesterday
Hiring Remotely in Canada
Remote
160K-190K Annually
Mid level
Artificial Intelligence • Big Data • Sports
AI insights + over 600 years of NFL expertise.
The Role
Design, build, and maintain reliable data pipelines and lakehouse assets for batch and streaming workloads. Implement ETL/ELT in Python and SQL, orchestration and monitoring (Airflow/Kubernetes), and retrieval/RAG/vector search pipelines for LLM and AI applications. Ensure data quality, lineage, versioning (Delta/Parquet/Iceberg), cost/performance optimization, and collaborate with MLOps and sports data teams.
Summary Generated by Built In

SumerSports is a leading football intelligence technology company that specializes in providing an innovative suite of products for football fans and NFL clubs. We are a collection of executives, engineers, data scientists, and visionaries from NFL clubs, technology startups, finance, and academia. 


Our data-driven platform empowers teams with insights and tools to make informed decisions within salary cap constraints. The platform also serves the NCAA, offering insights around the transfer portal and more.


What sets us apart is our unique blend of big tech talent, data scientists, and former NFL personnel, who have a combined 600+ years of NFL experience. Our domain knowledge is augmented by AI and machine learning technologies to create a unique view into many aspects of Football.

Position Summary


Our data engineering team is the foundation of this system, ensuring data is accurate, fast, and always available for our models and AI applications.


As a Data Engineer, you’ll design, build, and maintain the data pipelines that power our deep learning, video and LLM systems. You’ll work across ingestion, transformation, and orchestration layers — from real-time feeds to analytics-ready datasets. Your mission is to make data reliable, discoverable, and scalable for use by model training, analytics, and AI-driven products across multiple sports. You’ll collaborate closely with our MLOps, and Sports Data teams to ensure seamless integration between data and AI.


Responsibilities

  • Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Kubernetes.
  • Develop efficient ETL/ELT workflows in Python and SQL to support both batch and streaming workloads.
  • Partner with ML/AI teams to make datasets and tools discoverable and safe for autonomous agents, including evaluation and guardrails for AI-generated queries.
  • Develop retrieval pipelines (RAG, vector search) over structured stats and unstructured sources (scouting notes, video metadata) to power AI applications.
  • Model and maintain structured data assets (Delta, Parquet, Iceberg) for reliability, versioning, and lineage tracking.
  • Implement orchestration and monitoring: schedule jobs, track dependencies, and automate recovery from failures.
  • Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging.
  • Contribute to data platform evolution: evaluate tools, standardize best practices, and improve developer experience.
  • Support performance and cost optimization across compute, storage, and orchestration systems.


Qualifications

  • 3–8 years of experience as a Data Engineer or ETL Developer in a production environment.
  • Proficiency in Python and SQL; strong familiarity with Databricks, Spark, or equivalent big-data frameworks.
  • Experience with workflow orchestration tools such as Airflow, Dagster, Luigi or Prefect.
  • Deep understanding of data modeling, data warehousing, and distributed data processing.
  • Knowledge of modern data lakehouse architectures.
  • Familiarity with CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks.
  • Comfort working in a cross-functional environment with ML, product, and analytics teams.
  • Exposure to LLM-powered data tools: text-to-SQL, RAG, agent/tool interfaces (e.g. MCP), or natural-language analytics.
  • Previous work with cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Docker, Kubernetes).

Preferred

  • Previous experience with sports, telemetry, or sensor data pipelines.
  • Familiarity with streaming frameworks and event driven data processing (Kafka, Spark Structured Streaming, Flink).
  • General knowledge of American football, the NFL, and college football.
  • Background in data governance, lineage, and observability tools (Monte Carlo, Great Expectations, Unity Catalog, OpenLineage).
  • Experience designing semantic layers or metric definitions consumed by AI and BI tools.
  • Exposure to best practices in machine-learning model management and MLOps.

Benefits

  • Competitive Salary and Bonus Plan
  • Comprehensive health insurance plan
  • Retirement savings plan (401k) with company match
  • Remote working environment
  • A flexible, unlimited time off policy
  • Generous paid holiday schedule - 13 in total including Monday after the Super Bowl


SumerSports is committed to fair and equitable compensation practices.

Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.


The total compensation package for this position may also include annual performance bonus, benefits and/or other applicable incentive compensation plans.

Skills Required

  • 3-8 years of experience as a Data Engineer or ETL Developer in production
  • Proficiency in Python
  • Proficiency in SQL
  • Familiarity with Databricks or equivalent big-data frameworks (Spark)
  • Experience with workflow orchestration tools (Airflow, Dagster, Luigi, Prefect)
  • Deep understanding of data modeling, data warehousing, and distributed data processing
  • Knowledge of modern data lakehouse architectures
  • Familiarity with CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks
  • Exposure to LLM-powered data tools (text-to-SQL, RAG, agent/tool interfaces)
  • Previous work with cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Docker, Kubernetes)
  • Experience building and operating data pipelines for ingestion, cleaning, and transformation
  • Modeling and maintaining structured data assets (Delta/Parquet/Iceberg) for reliability and lineage
  • Experience with streaming frameworks and event-driven processing (Kafka, Spark Structured Streaming, Flink)
  • Previous experience with sports, telemetry, or sensor data pipelines
  • Background in data governance, lineage, and observability tools (Monte Carlo, Great Expectations, Unity Catalog, OpenLineage)
  • Experience designing semantic layers or metric definitions consumed by AI and BI tools
  • Exposure to best practices in machine-learning model management and MLOps
  • General knowledge of American football, the NFL, and college football
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palm Beach, FL
80 Employees
Year Founded: 2022

What We Do

SūmerSports is reshaping how professional and collegiate sports teams construct their rosters. Our top-tier data-driven platform empowers teams with invaluable insights and tools to make informed decisions while adhering to salary cap constraints. The platform will also service the NCAA, offering insights and recommendations around the transfer portal and much more. What sets us apart is our unique blend of big tech talent, data scientists, and former NFL personnel like John Idzik and Mike Mayock, who have a combined over 600 years of NFL experience. This unparalleled domain knowledge is seamlessly augmented by AI and machine learning technologies. Elevate your team's performance today with SūmerSports – where cutting-edge analytics converges with decades of real-world NFL experience, empowering you to outsmart the competition.

Similar Jobs

Movable Ink Logo Movable Ink

Data Engineer

Artificial Intelligence • Marketing Tech • Software
Easy Apply
Remote or Hybrid
Toronto, ON, CAN
600 Employees
130K-165K Annually

Samsara Logo Samsara

Data Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Canada
4000 Employees
119K-154K Annually

Movable Ink Logo Movable Ink

Software Engineer

Artificial Intelligence • Marketing Tech • Software
Easy Apply
Remote or Hybrid
Toronto, ON, CAN
600 Employees
130K-165K Annually
Remote
2 Locations
4172 Employees

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account