As our Senior Data Engineer, you'll own AQEMIA's data platform end to end — from ingestion through the pipeline to the trusted, model-ready datasets that power science, ML and analytics. What makes this role distinctive is the data itself: chemical structures, molecular conformations, physics- and ML-based predictions, and experimental results from CROs and partners.
A core part of the work is modelling these scientific entities well — establishing canonical identity, provenance and trustworthy lineage across heterogeneous, often messy sources — so scientists, and increasingly AI agents, can rely on them. You'll work at the intersection of software engineering, data infrastructure and scientific research, with real scope to shape architecture rather than just execute against it.
As AQEMIA moves toward more service and API-driven integration next year, you'll help make data fit for automation — expanding your impact from pipelines to the systems that consume them.
Responsibilities
- Own AQEMIA's Bronze → Silver → Gold data pipelines end to end, from ingestion through transformation and delivery, maintaining lineage and traceability as data volume and complexity grow.
- Model canonical scientific entities — compounds, structures, assays, predictions — establishing identity, provenance and trustworthy lineage across heterogeneous and often messy sources.
- Set and uphold data quality standards through monitoring, validation, testing and alerting across critical pipelines, strengthening governance and observability so datasets stay trusted and accessible.
- Partner with ML engineers, data scientists and researchers to build curated, model-ready datasets, translating scientific and business requirements into scalable data solutions.
- Drive data architecture and engineering best practices — data modeling, testing, documentation, orchestration and deployment — in collaboration with the Engineering Manager and Staff Data Engineer on roadmap execution.
- Build self-service capabilities and, looking ahead, APIs that make data fit for automation as AQEMIA moves toward more service-based integration.
- Uphold engineering quality through code reviews, and mentor junior engineers by sharing knowledge and best practices as a senior individual contributor.
Qualifications
- Deep experience in data engineering, with a track record of production systems that other teams depend on. We care about what you have built, not the year count.
- Strong software engineering skills. You code, you review code, and you hold a quality bar.
- Expert in Python and SQL. - Strong data modelling and relational database experience. You can defend an identity key.
- Production experience with a cloud warehouse (we use Snowflake) and a transformation framework (we use dbt).
- Experience with workflow orchestration (we use Airflow) and infrastructure-as-code (we use Terraform on AWS).
- Any STEM degree or equivalent experience.
Nice-to-have
- Experience with AWS.
- Experience with infrastructure-as-code (Terraform) and modern data warehousing (e.g. Snowflake, BigQuery, Redshift) and object storage.
- Experience in drug discovery, biotech, pharma or deeptech environments.
- Exposure to AI-driven or data-intensive workflows, or experience working across disciplines (e.g. biology ↔ ML ↔ chemistry).
- Experience implementing data governance, lineage and metadata management solutions.
- Track record of improving platform scalability, reliability and operational maturity.
Our recruitment process
- First discussion with our Talent Acquisition
- Hiring Manager’s interview: you’ll meet directly with your future manager
- Technical assessment of your skills in a deep-dive interview with the team
- VP interview to share wider team vision and align motivations
- Cultural fit interview with our co-founder and COO, Emmanuelle
- Final interview with our co-founder and CEO, Maximillien
Why Join Us?
Skills Required
- 7-10 years of experience in data engineering
- Strong software and data engineering skills
- Deep experience with data modeling and relational databases
- Strong proficiency in Python and SQL
- Experience building and maintaining production-grade data systems
- Hands-on experience with dbt, Airflow, or similar modern data stack tooling
- STEM degree or equivalent experience
- Experience with AWS
- Experience with infrastructure as code, including Terraform
- Experience with modern data warehousing such as Snowflake, BigQuery, or Redshift
- Experience with object storage
- Experience in drug discovery, biotech, pharma, or deeptech environments
- Exposure to AI-driven or data-intensive workflows
- Experience working across biology, machine learning, and chemistry disciplines
- Experience implementing data governance, lineage, and metadata management solutions
- Track record of improving platform scalability, reliability, and operational maturity
What We Do
AQEMIA is a next-gen pharmatech company generating one of the world's fastest-growing drug discovery pipeline. Our mission is to design fast innovative drug candidates for dozens of critical diseases, such as immuno-oncology. Our unique approach leverages quantum-inspired physics algorithms to power generative AI in designing novel drug candidates—without relying on experimental data. We already delivered several drug discovery successes within our internal pipeline and through collaborations with pharmaceutical companies. Our most advanced programs are currently in vivo optimization. We are growing and hiring! Check our career website: https://jobs.lever.co/aqemia.com Discover the roles and behind-the-scenes at AQEMIA on our Welcome To The Jungle page: https://www.welcometothejungle.com/fr/companies/aqemia





