SRE - DataPlatform

Reposted 6 Days Ago
Be an Early Applicant
Paris, Île-de-France, FRA
Hybrid
Mid level
eCommerce • Software
The Role
As an SRE, you will ensure the reliability and scalability of the data platform, implement SRE practices, support cloud migration, and collaborate with teams on performance optimization and disaster recovery.
Summary Generated by Built In

Join a transversal SRE community embedded in a product-oriented Data Platform team of 40–50 engineers, analysts, and data scientists across France and Spain. You'll drive the reliability and scalability of a next-generation Lakehouse platform — anchored on Trino, Iceberg, and on-prem object storage — while leading the transition from public cloud to a resilient hybrid/on-prem architecture.

🎯 TASKS 

    Platform Reliability & SRE foundations 
  • Own reliability of core data services: Trino, Iceberg, S3 / Ceph, Kafka, Kafka Connect, Schema Registry
  • Define and enforce SLIs/SLOs, error budgets, and on-call runbooks — solid SRE foundations are non-negotiable
  • Build full-stack observability with Prometheus and Grafana: metrics, dashboards, alerting pipelines, and anomaly detection
  • Manage and harden PostgreSQL clusters via Patroni for high-availability control-plane services
  • Kafka ecosystem — Connect & Schema governance
  • Operate and scale Kafka Connect clusters: connector lifecycle, offset management, dead-letter queues, and task rebalancing
  • Maintain the Schema Registry as the single source of truth for Avro/Protobuf/JSON schemas — enforce compatibility rules and schema evolution policies
  • Monitor consumer lag, connector throughput, and broker health via Prometheus JMX exporters and Grafana dashboards
  • Ensure end-to-end data contract integrity between producers and Iceberg/S3 consumers
  • Kubernetes, Kube-in-Kube & Crossplane
  • Operate production Kubernetes clusters (GKE/EKS + on-prem) — capacity planning, upgrades, PodDisruptionBudgets, resource quotas
  • Architect and manage Kube-in-Kube topologies to provide strong tenant isolation for data platform workloads — each team gets a dedicated virtual cluster without the overhead of a full physical cluster
  • Automate infrastructure and resource provisioning with Crossplane: define composite resources (XRDs) so data teams can self-serve Kafka topics, Trino namespaces, and S3 buckets through Kubernetes-native APIs
  • Maintain GitOps pipelines for platform deployment and configuration drift detection
  • Lakehouse architecture & cloud migration
  • Migrate from public cloud data warehouse to VeepeeCloud Iceberg-based lakehouse — managing coexistence, schema evolution, and time-travel
  • Architect resilient ingestion, transformation, and serving layers around Trino + S3
  • Optimize Trino query performance: memory limits, spilling, cost-based optimizer tuning
  • Agentic & developer enablement
  • Build agentic self-service tooling so data teams can provision Trino/Iceberg resources and Kafka Connect pipelines autonomously via Crossplane — reducing toil and ops bottlenecks
  • Develop FinOps dashboards (compute, storage, query cost) with Grafana and Prometheus-based cost exporters
  • Write clear technical documentation, runbooks, and internal ADRs
  • Multi-DC resilience & DRP
  • Design and implement multi-datacenter strategies across FR1 / NL1 — active-active and active-passive topologies
  • Leverage Fast Erasure Coding on object storage (Ceph/S3) to maximize durability with minimal replication overhead
  • Ensure data replication consistency across sites for Iceberg table metadata, Trino catalogs, and Schema Registry subjects
  • Lead DRP exercises: failover playbooks, RTO/RPO validation, postmortems

👉 MUST HAVE skills

    Must have
    • Strong experience with Kubernetes in production environments
    • Experience with Kube-in-Kube technologies (vCluster or similar)
    • Solid understanding of SRE principles (SLIs/SLOs, error budgets)
    • Experience with Prometheus and Grafana
    • Experience with Infrastructure as Code (Terraform or similar)
    • Experience with Crossplane
    • Familiarity with GitOps workflows
    • Experience with S3 and object storage technologies
    • Experience with PostgreSQL and Patroni
    • Experience with Kafka, Kafka Connect, and Schema Registry
    • Fluent in English
    •  
     
     

👉 NICE TO HAVE skills

  • Experience with multi-datacenter architectures (FR1/NL1)
  • Experience designing disaster recovery plans and failover playbooks
  • Experience with Fast Erasure Coding (Ceph/S3)
  • Experience with Trino, Iceberg, and Lakehouse technologies
  • Experience with Airflow
  • Experience building agentic self-service platforms
  • Knowledge of FinOps and cost optimization practices
  • Programming experience in Python, Java, or Go

✅ BENEFITS

  • Variable bonus
  • E-learning platform (self-education courses)
  • Meetups & conferences (local and international)
  • Flexible office — up to 2 days remote
  • International teams (France & Spain)

⚙️ RECRUITMENT PROCESS

  • 1️⃣ 30-minute HR Screen with a Veepeeᵀᵉᶜʰ  Recruiter

  • 2️⃣ General Technical exchange

  • 3️⃣ Technical exchange with the manager

  • 4️⃣ Team Interview

We are convinced that it is up to you to define the way you work, to develop yourself and to progress.

At Veepee we guarantee that you can just be yourself!

For the service of diversity and inclusion, Veepee is committed to reviewing all applications received on an equal basis.  

🔗COMPANY For more information about our ecosystem :  https://careers.veepee.com/en/home-page-en/ 

Skills Required

  • Strong experience with Kubernetes in production environments
  • Experience with distributed data systems
  • Solid understanding of SRE principles
  • Experience with Infrastructure as Code
  • Familiarity with GitOps workflows
  • Experience with observability tools
  • Comfortable working in cloud environments
  • Strong collaboration mindset
  • Fluent in English

Veepee Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Veepee and has not been reviewed or approved by Veepee.

  • Leave & Time Off Breadth Feedback suggests five weeks’ paid vacation, additional RTT accrual, and periodic extra closure weeks provide substantial time off for many roles. These provisions are highlighted as part of the standard package in core European locations.
  • Healthcare Strength Feedback suggests complementary health insurance via a mutuelle (often Alan) is a consistent element of the package. This anchors the basics of medical coverage in line with local norms.
  • Flexible Benefits Feedback suggests hybrid work is common, with regular remote-work options and occasional telework stipends. Meal support and public-transport reimbursement add everyday flexibility depending on site and role.

Veepee Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Paris
5,760 Employees
Year Founded: 2001

What We Do

The vente-privee group has consolidated its various European brands, together made up of 6000 employees, under one unified conglomerate: VEEPEE. This coalescence marks a new chapter in its European history. Present in 10 countries now, Veepee is taking a leading role in the European digital commerce landscape

Similar Jobs

Mondelēz International Logo Mondelēz International

Fp&a Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
Boulogne, Hauts-de-Seine, Île-de-France, FRA
90000 Employees

Mondelēz International Logo Mondelēz International

Chef de produit Senior Belin & Savory - CDI (H/F/X)

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
Clamart, Hauts-de-Seine, Île-de-France, FRA
90000 Employees

Tulip Logo Tulip

Forward Deployed Engineer - EMEA

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
27 Locations
310 Employees
70K-105K Annually

Datadog Logo Datadog

Manager I, Engineering - AI Platform - Evaluation & Annotation

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Paris, Île-de-France, FRA
6500 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account