Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
ResponsibilitiesData Processing at Scale: Build and maintain high-throughput image and video data pipelines—cleaning, filtering, augmenting, and transforming multimodal datasets
Data Strategy from Product Goals: Translate product objectives and training strategies into concrete data processing plans, defining required dataset characteristics, sourcing strategies, and preparation protocols for model consumption.
Iterative Quality Evaluation: Evaluate processed datasets against rigorous quality benchmarks, diagnose data-side failure modes, and iterate on processing strategies until datasets meet the high bar our foundation models demand.
Large-Scale Labeling Coordination: Coordinate and drive data labeling efforts with annotation teams, establishing labeling guidelines, reviewing outputs, and ensuring consistency across large-scale annotation campaigns.
You have strong Python programming skills and write clean, modular, production-grade data processing code.
You have hands-on experience with ML frameworks such as PyTorch—understanding model data requirements, tensor formats, and training data flow well enough to prepare data that researchers can consume directly.
You hold exceptionally high standards for data quality and can assess data reliability, accuracy, and coverage systematically.
You are familiar with Computer Vision domain concepts and understand how image and video data characteristics impact downstream generative model performance.
You communicate clearly, document your work thoroughly, and collaborate effectively with researchers, engineers, and labeling partners across time zones.
Production Quality, Agent Velocity: Your daily workflow runs through AI coding harnesses (e.g., AI agents/assistants), without sacrificing software engineering rigor. You review agent code diffs with the same scrutiny as a team member's PR, recognize AI code generation failure modes, and ship rapidly without introducing technical debt or "slop."
Experience with Computer Vision tasks related to image or video generation model training (e.g., diffusion models, autoregressive transformers, GANs).
Fluency with annotation formats such as COCO, PASCAL VOC, or custom labeling schemas.
Hands-on experience with data orchestration frameworks (e.g., Airflow, Dagster, Prefect, Luigi).
Experience with distributed data processing systems (e.g., Ray, Spark, Dask).
Skills Required
- Strong Python programming skills and ability to write clean, modular, production-grade data processing code
- Hands-on experience with ML frameworks such as PyTorch and understanding of model data requirements and tensor formats
- Exceptionally high standards for data quality and ability to assess data reliability, accuracy, and coverage systematically
- Familiarity with Computer Vision concepts and understanding how image/video characteristics impact generative model performance
- Clear communication, thorough documentation, and effective collaboration with researchers, engineers, and labeling partners across time zones
- Experience working with AI coding harnesses or agents and reviewing AI-generated code while maintaining software engineering rigor
- Experience with computer vision generation tasks (diffusion models, autoregressive transformers, GANs)
- Fluency with annotation formats such as COCO or PASCAL VOC or custom labeling schemas
- Hands-on experience with data orchestration frameworks (Airflow, Dagster, Prefect, Luigi)
- Experience with distributed data processing systems (Ray, Spark, Dask)
What We Do
Veeda AI is a small, fast-moving team of engineers and researchers building the next generation of multimodal foundation world models for Physical AI. Its work sits at the intersection of artificial intelligence, robotics, and embodied intelligence, with engineering roles involving high-throughput image and video data pipelines. The company aims to advance intelligent systems capable of operating in and understanding the physical world.








