Our growth is driving us to strengthen our GPU Cloud team to support our expanding infrastructure and key AI roadmap initiatives.
Your mission will be leading the Site Reliability Engineering (SRE) team in order to build, automate, and maintain a highly reliable, production-grade GPU cluster infrastructure powering our sovereign cloud.
We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.
You will be part of a team of 6 SREs within the GPU Cloud organization. The team focuses on critical AI and HPC infrastructure challenges, including automating key components of our stack and implementing support for modern GPU technologies.
Tasks
Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution
Design and implement automated solutions for server lifecycle management across GPU clusters
Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters
Plan, prioritize, and manage the technical development roadmap for the SRE team
Collaborate and coordinate closely with software engineering, product, and cross-functional teams across Scaleway
Handle recruitment and career management for team members
Maintain, scale, and optimize high-availability production systems under heavy load
Participate in on-call rotations to ensure production reliability and fast incident resolution
HARDSKILLS:
Strong experience managing engineering teams in high-constraint production environments
Proven expertise with Kubernetes container orchestration
Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)
Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)
SOFT SKILLS:
Strong engineering leadership and team management capabilities
Technical rigor and high attention to detail in production-critical environments
Ability to handle high-pressure operational situations and manage incident stress pragmatically
Excellent communication skills with the ability to convey challenging messages effectively
Collaborative mindset with a focus on empowering engineers rather than micromanaging
WHAT YOU WILL FIND AT SCALEWAY ++++
Hybrid work: We offer up to 3 days of remote work per week.
Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
🚀 Why join the Scaleway adventure?
✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
- Discovery call with HR
- Technical interview with the HPC team to understand your technical skills and approach to the role
- Manager interview to validate your expertise
- Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team
- HR interview and office visit to tour our offices and meet your future colleagues
Skills Required
- Proven experience managing engineering teams in production-critical environments
- Expertise with Kubernetes container orchestration
- Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
- Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
- Exposure to modern GPU hardware ecosystems (Nvidia, AMD)
- Experience with high-speed networking fabrics (InfiniBand, Spectrum-X, Tomahawk)
- Knowledge of distributed and high-performance storage solutions (Lustre, DDN, VAST)
- Ability to maintain, scale, and optimize high-availability production systems under heavy load
- Experience handling on-call rotations and incident management
- Strong engineering leadership, communication, and team development skills
What We Do
Welcome to Scaleway: Europe's empowering cloud provider We are the preferred cloud solution, empowering developers and businesses to seamlessly build, deploy, and scale applications across diverse infrastructures. Our global presence spans Paris, Amsterdam, and Warsaw, serving a thriving community of over 25,000 businesses, including pioneering European startups. Why Choose Us? Our comprehensive cloud ecosystem is distinguished by multi-AZ redundancy, delivering a robust foundation for uninterrupted operations. What sets us apart is not just our state-of-the-art data centers, powered entirely by renewable energy, but also our commitment to providing a smooth developer experience. Scaleway is the choice for those seeking native tools to navigate and manage multi-cloud architectures effortlessly. What We Offer: We offer fully managed solutions for bare metal, containerization, and serverless architectures. We believe in empowering our customers with choices: the choice to determine the location of their data, the choice to select the architecture that best suits their business needs, and the choice to scale responsibly. Our Advantages: Multi-AZ Redundancy: Ensure business continuity with our robust multi-AZ redundancy. Smooth Developer Experience: Streamline your development process with a user-friendly environment. Renewable Energy Data Centers: Contribute to a sustainable future by hosting your applications in data centers powered solely by renewable energy. Native Tools for Multi-Cloud: Seamlessly manage and optimize your multi-cloud architectures with our native tools. Choose us, Choose Innovation: We're not just a cloud provider; we are your partner in innovation. Choose us for the freedom to decide where your data resides, the flexibility to tailor your architecture to your business, and the responsibility to scale in an environmentally conscious manner.







