> Markdown version of [/jobs/ext/1416365-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1416365-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Mistral AI - **Location:** Paris, France - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Cloud Computing, Continuous Integration, DevOps, Distributed Systems, Fault Tolerance, Python (Programming Language), Reliability Engineering, Prometheus, Software Engineering, Web Services, Workflow Management Systems, Software Organization, Datadog, Data Logging, Scripting, System Availability, Grafana, Reliability of Systems, Cloudformation, Containerization, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Terraform, Docker, Elk Stack, Golang - **Published:** July 24, 2026 - **Apply:** https://fr.indeed.com/viewjob?jk=487cf89ea4cb321e ## About the Role * A Master's degree in Computer Science, Engineering, or a related field. * 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems. * Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations. * Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs. * Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes. * Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog. * Familiarity with infrastructure-as-code tools like Terraform or CloudFormation. * Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices. * Solid grasp of networking, security, and system administration concepts. * Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment. * Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus. ## Description As a Site Reliability Engineer (SRE) on the Platform team, you will shape the reliability, scalability, and performance of our platform and customer-facing applications. You'll work closely with software engineers and research teams to ensure our systems meet and exceed the expectations of both internal and external customers. This role balances day-to-day operations on production systems with long-term software engineering improvements. Your work will reduce operational toil, foster reliability, and ensure high availability for our web services, inference environments, and ML workloads. You'll enable seamless replication of work environments across multiple HPC clusters, directly impacting the stability and efficiency of our AI platform. What You Will Do * Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads. * Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters. * Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling. * Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance. * Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging. * Participate in on-call rotations to respond to incidents and perform root cause analysis. * Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform. * Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments. * Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams. * Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance. * Work with the security team to ensure infrastructure adheres to best practices and compliance requirements. * Document processes and procedures to ensure consistency and knowledge sharing across the team. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)