The role
As an SRE within our Core Engineering team, you’ll help operate and evolve Aqemia’s cloud platform end to end — from infrastructure-as-code and GitOps delivery to observability, security and FinOps.
The Core team owns the platform, and you’ll work to make it reliable, scalable, secure and easy to use. At Aqemia, nothing is deployed by hand and infrastructure changes flow through Git, so automation and reproducibility are central to how the team works.
The scale at Aqemia is different from a typical product company: bursty, large-scale parallel scientific computation, GPU fleets to plan and optimize, and some of the company's most valuable data to protect, all while keeping cost under control.
You’ll have the opportunity to contribute directly to architectural and tooling decisions, take ownership of meaningful parts of the platform, and work closely with the teams that depend on it every day. In a small Engineering organisation, improvements to the platform have a direct impact on how quickly our scientists and engineers can work.
As the platform evolves, from today's GitOps pipelines toward productionized MLOps and autonomous discovery workflows, you'll grow with it, taking on more scope rather than staying in a fixed lane.
Responsibilities
- Own AWS infrastructure end to end, built and maintained as code with OpenTofu and Terragrunt - no manual changes, no exceptions.
- Operate and evolve Kubernetes workloads via GitOps (ArgoCD, Helm, Kustomize), and drive adoption of standardized infrastructure patterns across teams.
- Contribute to the platform's reliability practice: own observability (metrics, logs, alerting), respond to incidents, and run blameless postmortems through to completed action items.
- Manage cloud cost as a shared responsibility - producing the monthly cost report, maintaining the cost allocation model, and partnering with teams to plan capacity ahead of large GPU compute campaigns.
- Set and enforce security posture across the platform: patching, vulnerability follow-up, and least-privileged access management.
- Build internal tooling and CI/CD pipelines that reduce friction for engineering, ML and scientific teams, and make the platform approachable to non-infrastructure users.
- Shape platform architecture and long-term strategy through technical reviews and sprint planning, sharing knowledge across infrastructure and DevOps topics.
Qualifications
- Strong platform/infrastructure engineering background, with 2-3+ years of experience post-degree.
- Deep hands-on expertise in AWS and infrastructure-as-code (Terraform/OpenTofu, Terragrunt).
- Strong production experience with Kubernetes and GitOps delivery (ArgoCD, Helm, Kustomize).
- Experience building and maintaining CI/CD pipelines (GitHub Actions or GitLab).
- Experience with cloud security practices and least-privileged access management.
Nice-to-have
- MLOps experience - training/inference pipelines, model lifecycle, workflow orchestrators.
- GPU capacity planning - autoscaling GPU fleets, spot strategies, quota management.
- Exposure to AI-driven or data-intensive workflows.
- Experience with another cloud provider beyond AWS (e.g. GCP).
Our recruitment process
- First discussion with our Talent Acquisition
- Hiring Manager’s interview: you’ll meet directly with your future manager
- Technical assessment of your skills in a deep-dive interview with the team
- Cultural fit interview with our co-founder and COO, Emmanuelle
- Final interview with our co-founder and CEO, Maximillien
Why Join Us?
Skills Required
- Strong platform or infrastructure engineering background with at least 2–3 years of post-degree experience
- Deep hands-on experience with AWS
- Deep hands-on experience with infrastructure as code, including Terraform or OpenTofu and Terragrunt
- Strong production experience with Kubernetes
- Experience with GitOps delivery using ArgoCD, Helm, or Kustomize
- Experience building and maintaining CI/CD pipelines with GitHub Actions or GitLab
- Experience with cloud security practices and least-privileged access management
- MLOps experience with training or inference pipelines, model lifecycle, or workflow orchestrators
- GPU capacity planning, including autoscaling GPU fleets, spot strategies, or quota management
- Exposure to AI-driven or data-intensive workflows
- Experience with a cloud provider beyond AWS, such as GCP
What We Do
AQEMIA is a next-gen pharmatech company generating one of the world's fastest-growing drug discovery pipeline. Our mission is to design fast innovative drug candidates for dozens of critical diseases, such as immuno-oncology. Our unique approach leverages quantum-inspired physics algorithms to power generative AI in designing novel drug candidates—without relying on experimental data. We already delivered several drug discovery successes within our internal pipeline and through collaborations with pharmaceutical companies. Our most advanced programs are currently in vivo optimization. We are growing and hiring! Check our career website: https://jobs.lever.co/aqemia.com Discover the roles and behind-the-scenes at AQEMIA on our Welcome To The Jungle page: https://www.welcometothejungle.com/fr/companies/aqemia








