Context & Vision
In 2026, writing code is no longer the primary bottleneck; managing its complexity and ensuring its reliability is. We are building a highly resilient advertising platform with a very lean, senior internal team, and we intend to keep it that way.
To achieve scale without the overhead of a large engineering department, we rely on an AI-first development paradigm and a philosophy borrowed from the best large-scale open-source projects. Our code is not public, but our governance is theirs: asynchronous communication, exhaustive written documentation, explicit rules, and uncompromising quality gates. It is the only way a small team of humans and AI agents ships serious systems without accumulating technical debt.
This role owns the ground the whole platform runs on: the clusters, the pipelines, and the guarantees that everything ships predictably.
The Role
You will own the platform and delivery infrastructure end to end — the clusters, the infrastructure-as-code, the delivery pipelines, and the operational guarantees behind them. In a lean team, reliability is not a separate department; it is a discipline you carry for everyone.
1. Infrastructure as Code. Own the cloud footprint through infrastructure-as-code and a GitOps workflow. Infrastructure is declared, reviewed and versioned like any other code — no click-ops, no undocumented state.
2. Delivery Pipelines.Own CI/CD. Builds are reproducible, deployments are predictable, and rollbacks are boring. You make shipping a non-event.
3. Reliability, Performance & DR. Own observability (metrics, logs, distributed traces), performance testing, and backup / disaster-recovery. You define the SLOs that matter for a real-time serving platform and you make them measurable.
4. Quality Gates & AI-First CI. Enforce quality at every passage point in the pipeline. Integrating AI into CI/CD — automated reviews, security and policy checks — is an open frontier here, and we expect you to study it and propose implementations. Everything you build is documented; if it is not written down, it does not exist.
The Tech Stack
- Orchestration: managed Kubernetes on Google Cloud Platform.
- IaC & GitOps: infrastructure-as-code and a GitOps workflow.
- CI/CD: modern delivery pipelines (and proposals to evolve them).
- Observability: metrics, logs, distributed tracing.
- Runtime context: Rust services, a React frontend, streaming, relational databases.
- Cloud:Google Cloud Platform.
Profile & Requirements
We are looking for a platform engineer who treats infrastructure as a product, owns reliability for the whole team, and is comfortable in a lean, high-quality, AI-first environment. This role is not suited to someone who wants to run a ticket queue inside a large ops team.
Essential Experience
- 4+ years in DevOps / platform / SRE roles, running production Kubernetes** for real workloads (GCP preferred).
- Infrastructure-as-code and a GitOps workflow in production; infrastructure declared and reviewed as code.
- Solid CI/CD ownership and a real observability practice (metrics, logs, distributed tracing)
- Experience defining and defending SLOs, performance and disaster-recovery for latency-sensitive systems.
Core Competencies & Mindset
- AI Development Lifecycle: comfort integrating AI into the delivery workflow, and a point of view on automating quality and security gates in CI.
- Uncompromising Reliability: a maniacal focus on predictability, recoverability and security. Surprises in production are the enemy.
- Written & Asynchronous: exceptional written communication; runbooks and ADRs are part of the job, not a favour. Effective in a distributed, async environment (Paris timezone +/- 3h).
- Ownership:** you carry reliability for the whole team and raise risks before they become incidents.
Skills Required
- 4+ years of experience in DevOps, platform engineering, or SRE roles
- Production Kubernetes experience running real workloads
- Production experience with infrastructure as code and GitOps workflows
- Solid CI/CD ownership experience
- Experience implementing observability using metrics, logs, and distributed tracing
- Experience defining and defending SLOs
- Experience with performance testing and disaster recovery for latency-sensitive systems
- Comfort integrating AI into delivery workflows and automating quality and security gates in CI
- Exceptional written communication and experience creating runbooks and ADRs
- Ability to work effectively in a distributed, asynchronous environment
What We Do
Yassir is the leading super App in the Maghreb region set to changing the way daily services are provided. It currently operates in 45 cities across Algeria, Morocco and Tunisia with recent expansions into France, Canada and Sub-Saharan Africa. It is backed (~$200M in funding) by VCs from Silicon Valley, Europe and other parts of the world. We offer on-demand services such as ride-hailing and last-mile delivery. Building on this infrastructure, we are now introducing financial services to help our users pay, save and borrow digitally. Helping usher the continent into a digital economy era. We’re not just about serving people - we’re about creating a marketplace to bring people what they need while infusing social values






