Data Engineer - Foundational

Posted Yesterday
Be an Early Applicant
Paris, Île-de-France
In-Office
200K-200K Annually
Senior level
Artificial Intelligence • Computer Vision • Machine Learning • Robotics • Defense • Manufacturing
Building the Future of Autonomous Warfare. With Speed and Intelligence.
The Role
As a Data Engineer, build and optimize high-performance data infrastructures for deep learning by managing large volumes of raw video data, ensuring efficient processing pipelines and data quality.
Summary Generated by Built In
About Us

Harmattan AI is a next-generation defense prime building autonomous and scalable defense systems. Following the close of a $200M Series B, valuing the company at $1.4 billion, we are expanding our teams and capabilities to deliver mission-critical systems to allied forces.

Our work is guided by clear values: building technologies with real-world impact, pursuing excellence in everything we do, setting ambitious goals, and taking on the hardest technical challenges. We operate in a demanding environment where rigor, ownership, and execution are expected.

About the Role

As a Data Engineer on the Foundational team, you will serve as the "plumber" for deep learning, building the massive, high-performance data infrastructure required to power our foundational models. Based in Paris, you will manage terabytes—and eventually petabytes—of raw, unstructured, and noisy video data (EO and IR). Your mission is to ensure our ML engineers spend their time designing architectures, not waiting for data loaders or wrangling corrupted files.


Responsibilities
  • Multi-Modal Ingestion Pipeline: Build ETL/ELT pipelines to extract, decode, and store raw Electro-Optical (EO) and Infrared (IR) video from field logs into highly optimised formats like WebDataset, TFRecords, or Parquet.

  • Sensor Synchronisation & Alignment: Develop algorithms to programmatically synchronise EO and IR frames temporally and spatially to provide paired inputs for model training.

  • High-Throughput Data Loading: Architect storage-to-GPU pipelines to ensure multi-node training clusters maintain >90% GPU utilisation without I/O bottlenecks.

  • Distributed Processing: Write and optimise distributed data processing jobs using tools like Apache Spark, Ray, or Apache Beam to process thousands of hours of tactical video logs.

  • Data Quality & Versioning: Implement automated quality checks to filter corrupted or blank frames and maintain 100% reproducible training runs through robust versioning and lineage tracking.

  • Infrastructure Evaluation: Assess and implement advanced storage solutions (e.g., MinIO, S3 tiering) to manage growing datasets while optimising for cost and latency.


Candidate Requirements
  • Educational Background: A BS or MS in Computer Science, Software Engineering, or Distributed Systems is highly preferred. Deep knowledge of operating systems, networking, and parallel computing is essential.

  • Technical Experience: 5-6+ years of experience building and maintaining terabyte-scale pipelines for unstructured data (video, images, or point clouds).

  • Performance Optimisation: Proven track record of maximising multi-node GPU utilisation and optimising data loaders for frameworks like PyTorch or JAX.

  • Tooling Expertise: Strong command of distributed computing tools (Spark, Ray, Beam) and ML data versioning tools (DVC, Apache Iceberg, or Pachyderm).

  • Adaptability & Ownership: A systems-thinker who thrives in a fast-paced startup environment and views messy data as an engineering problem to be solved via automation.

  • Commitment: 100% dedication to Harmattan AI’s mission of providing a defensive edge to allied nations through ethical, high-impact technology

We look forward to hearing how you can help shape the future of autonomous defense systems at Harmattan AI.

Top Skills

Apache Beam
Apache Iceberg
Spark
Dvc
Elt
ETL
Jax
Pachyderm
Parquet
PyTorch
Ray
Tfrecords
Webdataset
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Paris, Île-de-France
131 Employees

What We Do

Harmattan AI is rising as a next-generation defense prime, building the future of autonomous warfare. We leverage AI-driven autonomy, real-time intelligence, and conflict-ready production to deliver attritable systems and autonomous mission management software. Designed for the real-world needs of warfighters, our solutions enable faster deployment, sharper decision-making, and battlefield dominance.

Similar Jobs

Cloudflare Logo Cloudflare

Senior Solutions Architect

Cloud • Information Technology • Security • Software • Cybersecurity
Hybrid
5 Locations
4400 Employees

InterSystems Logo InterSystems

Business Development Manager

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Easy Apply
In-Office
Courbevoie, Hauts-de-Seine, Île-de-France, FRA
2407 Employees

ServiceNow Logo ServiceNow

EMEA - Senior Business Strategy Manager

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Issy-les-Moulineaux, Hauts-de-Seine, Île-de-France, FRA
28000 Employees

Datadog Logo Datadog

Senior Security Engineer

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Paris, Île-de-France, FRA
6500 Employees

Similar Companies Hiring

Idler Thumbnail
Artificial Intelligence
San Francisco, California
6 Employees
Fairly Even Thumbnail
Software • Sales • Robotics • Other • Hospitality • Hardware
New York, NY
Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account