- Design the lakehouse layer over our existing S3 and Parquet footprint, including table format selection (Iceberg or Delta), partitioning strategy, schema evolution, and a searchable catalogue.
- Decide what stays as raw imagery and what becomes a managed table, and document why.
- Stand up orchestration from scratch.
- Evaluate and select the tooling (Airflow, Dagster, Prefect, or equivalent), then build batch pipelines that are idempotent, backfillable, and observable.
- Build reliable ingestion for drone imagery and telemetry captured in the field, often over poor connectivity.
- Handle validation, deduplication, and integrity checks close to capture.
- Work directly with AI engineers to turn raw captures into curated, versioned training and evaluation datasets.
- Own dataset lineage and versioning so that any model in MLflow can be traced back to the exact data that produced it, two years later. Support annotation workflows and our vector database.
- Define data contracts, automated tests, and freshness and anomaly monitoring.
- Own lineage from raw capture through to the datasets and features that models consume. Establish the data quality standards the team codes against.
- Own S3 storage class lifecycle, file compaction, and Athena scan cost.
- At our scale, storage and query layout decisions are the primary cost lever, and we expect you to treat that as an engineering problem rather than a finance one.
- As the platform matures, extend it to serve business and product analytics: a modelled warehouse layer, transformation tooling such as dbt, and a semantic layer for reporting.
- This is a later phase of the role, not a day-one responsibility, and you will help decide when it becomes the priority.
- Write production-grade Python.
- Use infrastructure as code, containerization, CI/CD, and version control as defaults rather than afterthoughts.
- Raise the bar across the AI team through code review and by setting the data engineering standards other engineers build on.
- Communicate design decisions, tradeoffs, and progress clearly to engineers and to non-technical partners.
- 4+ years building and operating production data platforms, with clear ownership of systems you designed rather than only maintained.
- Strong Python, with solid software engineering fundamentals: testing, code structure, version control, and code review.
- Deep hands-on AWS experience, particularly S3, Athena or equivalent query engines, IAM, and cost management at scale.
- Production experience with a workflow orchestrator (Airflow, Dagster, Prefect, Step Functions, or similar).
- Advanced SQL and demonstrable data modeling judgment.
- Experience with columnar formats and open table formats: Parquet plus Iceberg, Delta, or Hudi.
- Track record of designing for reproducibility, including data versioning, lineage, and backfill correctness.
- Comfort operating without an existing platform to lean on and the judgment to sequence what gets built first.
- Geospatial and raster data experience: GeoTIFF, cloud-optimized GeoTIFF, GDAL, tiling, and coordinate reference systems.
- Experience supporting computer vision or ML teams, including training dataset curation, annotation pipelines, or feature stores.
- Familiarity with MLflow, DVC, LakeFS, or comparable experiment and data versioning tooling.
- Distributed processing experience: Spark, Ray, or Dask.
- Analytics engineering exposure: dbt, dimensional modelling, BI tooling.
- Infrastructure as code: Terraform, CDK, or Pulumi.
- Kubernetes.
- Agriculture, remote sensing, robotics, or another domain with large sensor-derived datasets.
- Bachelor's or master's degree in computer science, computer engineering, software engineering, and data science.
Skills Required
- 4+ years building and operating production data platforms, with ownership of systems designed
- Strong Python and software engineering fundamentals, including testing, code structure, version control, and code review
- Deep hands-on AWS experience, particularly S3, Athena or equivalent query engines, IAM, and cost management at scale
- Production experience with a workflow orchestrator such as Airflow, Dagster, Prefect, or Step Functions
- Advanced SQL and demonstrable data modeling judgment
- Experience with Parquet and open table formats such as Iceberg, Delta, or Hudi
- Experience designing for reproducibility, including data versioning, lineage, and backfill correctness
- Ability to operate independently without an existing data platform and prioritize what to build first
- Bachelor's or master's degree in computer science, computer engineering, software engineering, or data science
- Geospatial and raster data experience, including GeoTIFF, Cloud-Optimized GeoTIFF, GDAL, tiling, and coordinate reference systems
- Experience supporting computer vision or machine learning teams, including training dataset curation, annotation pipelines, or feature stores
- Familiarity with MLflow, DVC, LakeFS, or comparable experiment and data versioning tools
- Distributed processing experience with Spark, Ray, or Dask
- Analytics engineering exposure, including dbt, dimensional modeling, or BI tooling
- Infrastructure-as-code experience with Terraform, CDK, or Pulumi
- Kubernetes experience
- Experience in agriculture, remote sensing, robotics, or another sensor-data domain
What We Do
At Precision AI we are on a mission to accelerate artificial intelligence based farming practices to create healthier, happier, and more profitable farms. By leveraging our advanced drones and custom-built AI technology, we can take crop production decisions from a whole field to an individual plant level. This type of decision-making transforms an industry that has been reliant on larger and broader technology for decades. The outcome of our solutions is integrated into the agricultural technology of today and helps craft the machines of tomorrow that will feed the world. Precision AI was founded in 2017 with headquarters in Regina, Saskatchewan. We are scaling rapidly with an elite global team solving the agriculture challenges of farms around the world.









