About This Role
Modern ML systems improve through their data flywheel: production experience generates new data, that data is processed and understood, valuable examples are identified and labeled, datasets are created, models are retrained and evaluated, and the resulting models return to production.
We are looking for an ML Engineer focused on the Data Flywheel to help build and automate this lifecycle for autonomous robots operating in complex real-world environments.
This role sits between ML engineering and data engineering. You will work closely with perception/autonomy researchers, labeling, Data Platform, and ML Platform engineers to turn large volumes of robot experience into high-quality training data.
What You'll Get to Do
- Build and maintain pipelines that transform robot data into training-ready ML datasets.
- Automate data processing, filtering, selection, transformation, labeling, and dataset-generation workflows.
- Develop scalable batch and offline inference pipelines for auto-labeling and data mining.
- Build systems for selecting useful or difficult examples from large volumes of robot data.
- Integrate manual labeling and model-assisted/automated labeling into repeatable data workflows.
- Implement data-quality checks and dataset validation.
- Build reproducible and versioned dataset-generation pipelines.
- Improve incremental dataset generation so new robot experience can efficiently feed future training cycles.
- Partner with researchers to translate experimental data-processing logic into reliable production pipelines.
- Measure and improve the efficiency of the loop from new robot experience to usable training data.
What We're Looking For
- Strong Python and software engineering skills.
- Experience building ML/data-processing pipelines.
- Understanding of ML dataset construction and training workflows.
- Experience processing large datasets using distributed or parallel computing.
- Experience with object storage, Parquet or similar formats, workflow systems, and cloud compute.
- Ability to bridge exploratory ML code and reliable production systems.
- Strong understanding of data quality and reproducibility.
Nice to Have
- Computer vision, perception, robotics, or multimodal ML experience.
- Experience with auto-labeling, active learning, hard-example mining, or data selection.
- Experience processing images, video, point clouds, LiDAR, or other sensor data.
- Ray/Spark or similar distributed-processing experience.
Skills Required
- Strong Python and software engineering skills
- Experience building ML or data-processing pipelines
- Understanding of ML dataset construction and training workflows
- Experience processing large datasets using distributed or parallel computing
- Experience with object storage, Parquet or similar formats, workflow systems, and cloud compute
- Ability to bridge exploratory ML code and reliable production systems
- Strong understanding of data quality and reproducibility
- Computer vision, perception, robotics, or multimodal ML experience
- Experience with auto-labeling, active learning, hard-example mining, or data selection
- Experience processing images, video, point clouds, LiDAR, or other sensor data
- Ray, Spark, or similar distributed-processing experience
What We Do
FieldAI is pioneering the development of a field-proven, hardware agnostic brain technology that enables many different types of robots to operate autonomously in hazardous, offroad, and potentially harsh industrial settings – all without GPS, maps, or any pre-programmed routes.








