Machine Learning Engineer - Data Curation

Posted Yesterday
Be an Early Applicant
Toronto, ON, CAN
Hybrid
Mid level
Artificial Intelligence • Computer Vision • Machine Learning • Robotics
The Role
Build and maintain high-throughput image/video data pipelines, define dataset strategies from product goals, evaluate data quality, and coordinate large-scale labeling to prepare multimodal datasets for foundation model training.
Summary Generated by Built In
About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Data Processing at Scale: Build and maintain high-throughput image and video data pipelines—cleaning, filtering, augmenting, and transforming multimodal datasets

  • Data Strategy from Product Goals: Translate product objectives and training strategies into concrete data processing plans, defining required dataset characteristics, sourcing strategies, and preparation protocols for model consumption.

  • Iterative Quality Evaluation: Evaluate processed datasets against rigorous quality benchmarks, diagnose data-side failure modes, and iterate on processing strategies until datasets meet the high bar our foundation models demand.

  • Large-Scale Labeling Coordination: Coordinate and drive data labeling efforts with annotation teams, establishing labeling guidelines, reviewing outputs, and ensuring consistency across large-scale annotation campaigns.

Requirements
  • You have strong Python programming skills and write clean, modular, production-grade data processing code.

  • You have hands-on experience with ML frameworks such as PyTorch—understanding model data requirements, tensor formats, and training data flow well enough to prepare data that researchers can consume directly.

  • You hold exceptionally high standards for data quality and can assess data reliability, accuracy, and coverage systematically.

  • You are familiar with Computer Vision domain concepts and understand how image and video data characteristics impact downstream generative model performance.

  • You communicate clearly, document your work thoroughly, and collaborate effectively with researchers, engineers, and labeling partners across time zones.

  • Production Quality, Agent Velocity: Your daily workflow runs through AI coding harnesses (e.g., AI agents/assistants), without sacrificing software engineering rigor. You review agent code diffs with the same scrutiny as a team member's PR, recognize AI code generation failure modes, and ship rapidly without introducing technical debt or "slop."

Nice to Have
  • Experience with Computer Vision tasks related to image or video generation model training (e.g., diffusion models, autoregressive transformers, GANs).

  • Fluency with annotation formats such as COCO, PASCAL VOC, or custom labeling schemas.

  • Hands-on experience with data orchestration frameworks (e.g., Airflow, Dagster, Prefect, Luigi).

  • Experience with distributed data processing systems (e.g., Ray, Spark, Dask).

Skills Required

  • Strong Python programming skills and ability to write clean, modular, production-grade data processing code
  • Hands-on experience with ML frameworks such as PyTorch and understanding of model data requirements and tensor formats
  • Exceptionally high standards for data quality and ability to assess data reliability, accuracy, and coverage systematically
  • Familiarity with Computer Vision concepts and understanding how image/video characteristics impact generative model performance
  • Clear communication, thorough documentation, and effective collaboration with researchers, engineers, and labeling partners across time zones
  • Experience working with AI coding harnesses or agents and reviewing AI-generated code while maintaining software engineering rigor
  • Experience with computer vision generation tasks (diffusion models, autoregressive transformers, GANs)
  • Fluency with annotation formats such as COCO or PASCAL VOC or custom labeling schemas
  • Hands-on experience with data orchestration frameworks (Airflow, Dagster, Prefect, Luigi)
  • Experience with distributed data processing systems (Ray, Spark, Dask)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Veeda AI is a small, fast-moving team of engineers and researchers building the next generation of multimodal foundation world models for Physical AI. Its work sits at the intersection of artificial intelligence, robotics, and embodied intelligence, with engineering roles involving high-throughput image and video data pipelines. The company aims to advance intelligent systems capable of operating in and understanding the physical world.

Similar Jobs

Samsara Logo Samsara

Mid-market Account Executive

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Toronto, ON, CAN
4000 Employees
159K-175K Annually

Samsara Logo Samsara

Mid-market Account Executive

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Ottawa, ON, CAN
4000 Employees
159K-175K Annually

TransUnion Logo TransUnion

Director, Credit Risk Solutions Consulting

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
2 Locations
13000 Employees
174K-245K Annually

DraftKings Logo DraftKings

Senior Associate Delivery Manager

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Remote or Hybrid
Canada
6400 Employees
82K-102K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account