> Markdown version of [/jobs/ext/1943277-sre-engineering-manager-gpu-cloud-h-f](https://www.wearedevelopers.com/jobs/ext/1943277-sre-engineering-manager-gpu-cloud-h-f). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sre Engineering Manager - Gpu Cloud H/F - **Company:** Bienvenue Chez Scaleway - **Location:** Paris, France (Remote available) - **Contract:** Permanent contract - **Skills:** Proxmox, Artificial Intelligence, Cloud Computing, Computer Clusters, InfiniBand, Reliability Engineering, Prometheus, Software Engineering, Virtualization Technology, Data Logging, Grafana, Containerization - **Published:** August 6, 2026 - **Apply:** https://www.hellowork.com/fr-fr/emplois/82080344.html ## About the Role Strong experience managing engineering teams in high-constraint production environments - Proven expertise with Kubernetes container orchestration - Direct experience with cluster management and virtualization tools (Proxmox, Warewulf) - Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana) - Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk) - Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST) SOFT SKILLS: - Strong engineering leadership and team management capabilities - Technical rigor and high attention to detail in production-critical environments - Ability to handle high-pressure operational situations and manage incident stress pragmatically - Excellent communication skills with the ability to convey challenging messages effectively - Collaborative mindset with a focus on empowering engineers rather than micromanaging ## Description Our growth is driving us to strengthen our GPU Cloud team to support our expanding infrastructure and key AI roadmap initiatives. Your mission will be leading the Site Reliability Engineering (SRE) team in order to build, automate, and maintain a highly reliable, production-grade GPU cluster infrastructure powering our sovereign cloud. YOUR FUTURE TEAM We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together. You will be part of a team of 6 SREs within the GPU Cloud organization. The team focuses on critical AI and HPC infrastructure challenges, including automating key components of our stack and implementing support for modern GPU technologies. YOUR DAILY ROUTINE Tasks - Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution - Design and implement automated solutions for server lifecycle management across GPU clusters - Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters - Plan, prioritize, and manage the technical development roadmap for the SRE team - Collaborate and coordinate closely with software engineering, product, and cross-functional teams across Scaleway - Handle recruitment and career management for team members - Maintain, scale, and optimize high-availability production systems under heavy load - Participate in on-call rotations to ensure production reliability and fast incident resolution ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [Best Companies to Work For in Paris: Top 25 Companies in 2023 ](https://www.wearedevelopers.com/magazine/190-best-companies-to-work-for-in-paris-top-25-companies-in-2023) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Companies to Work For in France: Top 25 Companies in 2023 ](https://www.wearedevelopers.com/magazine/189-best-companies-to-work-for-in-france-top-25-companies-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Welcome to Switzerland](https://www.wearedevelopers.com/magazine/4-welcome-to-switzerland) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)