About You:
You think data infrastructure is a craft. You have built pipelines that ran unattended for months and you have felt the specific satisfaction of a dataset that is exactly what it claims to be. You'd rather prevent a data problem than debug a model trained on one.
You will own the machinery that turns flight and simulation output into trustworthy, versioned, traceable training data — and turns deployed behavior back into the next training set.
Responsibilities:
- Build and operate ingestion pipelines for flight test, simulation and operational data, including format normalization, time alignment across sources, and quality validation at the gate.
- Implement dataset curation and versioning with full lineage: what went into a dataset, from where, processed how, and by which version of which tool.
- Build automated mining and triage for the rare and interesting — anomalies, disagreements between learned and rule-based systems, near-boundary events — so the flywheel prioritizes signal over volume.
- Own labeling workflows and tooling, including quality control and inter-annotator agreement where human labeling is involved.
- Close the loop: instrument deployed-system behavior so that operational signal reliably reaches the next training cycle without manual shepherding.
- Work with Flight Test & Operations on logging and instrumentation requirements, and help resolve access constraints where data is captured by third-party systems.
- Monitor pipeline health and dataset coverage; alert on drift, gaps and silent breakage.
Qualifications:
- Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or a related subject.
- 3+ years in data engineering, with production ownership of pipelines feeding ML training.
- Strong Python and SQL; solid grasp of data modeling, storage formats and processing frameworks.
- Experience with dataset versioning and lineage tooling.
- Comfort with high-volume time-series and multimodal sensor data — telemetry, video, audio, structured logs.
- Care about correctness. In this role a quiet data bug becomes a model behavior nobody can explain.
Nice to Have:
- Robotics, autonomous vehicle or aerospace data experience.
- Familiarity with flight data formats and avionics bus data.
- Experience building internal tools that engineers actually chose to use.
- Data governance or provenance work in a regulated environment.
Skills Required
- 3+ years of experience in data engineering
- Production ownership of pipelines feeding machine learning training
- Strong Python skills
- Strong SQL skills
- Understanding of data modeling, storage formats, and processing frameworks
- Experience with dataset versioning and lineage tooling
- Experience with high-volume time-series and multimodal sensor data, including telemetry, video, audio, and structured logs
- Robotics, autonomous vehicle, or aerospace data experience
- Familiarity with flight data formats and avionics bus data
- Experience building internal engineering tools
- Data governance or provenance experience in a regulated environment
What We Do
All of the sky, none of the limits. Merlin Labs is building the autonomous infrastructure for the sky above us, enabling goods, and eventually people, to fly without pilots. The Merlin Labs team is building the definitive autonomy system for all things that fly. We’re creating sophisticated software and hardware that fulfill the functions of a human pilot.









