About the Role
We are building the data foundation that powers the full machine learning lifecycle for autonomous robotics. Our robots generate large-scale, multimodal datasets across real-world deployments, and turning that raw experience into reliable, discoverable, high-quality ML data is a core part of improving our autonomy systems.
As a Data Platform Engineer, you will help design and build the platform that manages data from ingestion through processing, validation, labeling, dataset generation, training, and evaluation.
This is not a traditional analytics data engineering role. You will work closely with ML engineers, researchers, labeling teams, robotics engineers, and infrastructure engineers to build scalable systems for robotics and ML data.
What You'll Do
- Design and build scalable data architecture for large-scale multimodal robotics and ML datasets.
- Build abstractions and services for ingestion, processing, datasets, metadata, lineage, and data quality.
- Develop reliable pipelines for transforming raw robot data into versioned, ML-ready data products.
- Define data models, schemas, contracts, and lifecycle states across data-processing workflows.
- Build systems for tracking provenance and lineage across raw data, derived artifacts, labels, datasets, and downstream ML workloads.
- Develop automated validation and data-quality frameworks that detect incomplete, corrupted, or unusable data early.
- Design for incremental processing, reprocessing, backfills, and versioned transformations.
- Improve observability and failure diagnosis across complex data workflows.
- Partner with ML and robotics teams to understand domain-specific data requirements and turn recurring patterns into reusable platform capabilities.
- Work closely with infrastructure/platform teams on storage, compute, orchestration, reliability, and scalability.
What We're Looking For
- Strong experience building production data platforms or large-scale data-processing systems.
- Strong software engineering skills, preferably Python and/or C++/Java/Go.
- Experience with distributed data processing and workflow orchestration.
- Experience with data lakes/lakehouses, object storage, metadata systems, schemas, and data versioning.
- Strong understanding of data quality, lineage, reproducibility, and reliable pipeline design.
- Experience with technologies such as S3, Airflow/Dagster, Spark/Ray, Kubernetes, Parquet, or similar systems.
- Ability to work across ambiguous organizational and technical boundaries.
- Strong systems-design and engineering judgment.
Nice to Have
- Experience with ML datasets or ML infrastructure.
- Robotics, autonomous vehicles, sensor, video, image, LiDAR, or other multimodal data.
- Experience building internal developer/platform products.
- Experience operating pipelines at TB/PB scale.
Skills Required
- Strong experience building production data platforms or large-scale data-processing systems
- Strong software engineering skills in Python and/or C++, Java, or Go
- Experience with distributed data processing and workflow orchestration
- Experience with data lakes or lakehouses, object storage, metadata systems, schemas, and data versioning
- Strong understanding of data quality, lineage, reproducibility, and reliable pipeline design
- Experience with systems such as S3, Airflow or Dagster, Spark or Ray, Kubernetes, and Parquet
- Ability to work across ambiguous organizational and technical boundaries
- Strong systems-design and engineering judgment
- Experience with ML datasets or ML infrastructure
- Experience with robotics, autonomous vehicles, sensor, video, image, LiDAR, or other multimodal data
- Experience building internal developer or platform products
- Experience operating pipelines at terabyte or petabyte scale
What We Do
FieldAI is pioneering the development of a field-proven, hardware agnostic brain technology that enables many different types of robots to operate autonomously in hazardous, offroad, and potentially harsh industrial settings – all without GPS, maps, or any pre-programmed routes.







