Applied AI Engineer, Site Reliability Engineer - EMEA

Mistral AI
Paris, France
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cyber Security Computer Networks DevOps Python (Programming Language) Linux Kernel Systems Development Life Cycle Role-Based Access Control Reliability Engineering Site Reliability Engineering Practices
+13 more
Ansible Prometheus Google Cloud Large Language Models Grafana Software Security Multi-Cloud Build Management Kubernetes Deployment Automation CIS Benchmarks Terraform Golang

Job description

Experteer Overview As a founding member of the Applied AI SRE sub-team, you will design and operate a fleet-focused reliability framework for Mistral’s AI solutions. You’ll scale SRE practices across Mistral-hosted and customer-hosted deployments, driving observability, runbooks, and security guardrails. You’ll own on-call incidents, post-mortems, and CVE response while partnering with product and security teams. This is a chance to shape how we deliver robust, scalable AI for enterprise clients at scale. You’ll work across a fast-paced environment with a strong emphasis on impact, ownership, and collaboration. Pay / Benefits * Design and build a fleet-wide reliability framework (SLOs, observability, runbooks) * Run Tier-1 customer environments, maintain SLO compliance, on-call and incident response * Productize deployment, security baselines, and scale of Applied AI solutions * Own security operations for customer deployments, including CVE response and supply-chain integrity * Collaborate with Technical Support as L3 escalation, perform blameless post-mortems * Automate provisioning and enforce secure-config baselines * Lead security and operational excellence across Mistral-hosted and customer-hosted fleets Tasks * Fluent in English * 5+ years in SRE, Production Engineering, or DevOps with tooling shipping track record * Strong multi-tenant Kubernetes fluency (namespaces, network policies, RBAC, admission control, scale) * On-call discipline, incident response, blameless post-mortem culture * Observability stack in production (Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Signoz) * Infrastructure as code (Terraform, Ansible or equivalents) * Proficiency in Python and/or Golang for tooling and automation * Security mindset: secure-SDLC, CVE response, supply-chain integrity as reliability properties * Strong written communication (runbooks, post-mortems, customer incident comms) * Ability to operate autonomously in ambiguous, fast-paced environments * Solid Linux internals, networking, distributed-systems fundamentals * Nice to have: cloud/app security (AppSec, K8s security, SBOM, cosign, SLSA) * Nice to have: production experience with LLM/model-serving stacks * Nice to have: multi-cloud or on-prem hybrid environments (AWS, GCP, Azure, sovereign clouds) * Nice to have: open-source contributions in SRE/observability/security tooling Key requirements * healthcare coverage * parental leave * retirement plans * relocation support * wellness programs * meal and transportation allowances

Requirements

Collaborate with Technical Support as L3 escalation, perform blameless post-mortems * Automate provisioning and enforce secure-config baselines * Lead security and operational excellence across Mistral-hosted and customer-hosted fleets Tasks * Fluent in English * 5+ years in SRE, Production Engineering, or DevOps with tooling shipping track record * Strong multi-tenant Kubernetes fluency (namespaces, network policies, RBAC, admission control, scale) * On-call discipline, incident response, blameless post-mortem culture * Observability stack in production (Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Signoz) * Infrastructure as code (Terraform, Ansible or equivalents) * Proficiency in Python and/or Golang for tooling and automation * Security mindset: secure-SDLC, CVE response, supply-chain integrity as reliability properties * Strong written communication (runbooks, post-mortems, customer incident comms) * Ability to operate autonomously in ambiguous, fast-paced environments * Solid Linux internals, networking, distributed-systems fundamentals * Nice to have: cloud/app security (AppSec, K8s security, SBOM, cosign, SLSA) * Nice to have: production experience with LLM/model-serving stacks * Nice to have: multi-cloud or on-prem hybrid environments (AWS, GCP, Azure, sovereign clouds) * Nice to have: open-source contributions in SRE/observability/security tooling Key requirements * healthcare coverage * parental leave * retirement plans * relocation support * wellness programs * meal and transportation allowances

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eu.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

6:41 min

Outsourcing personal organization and schedules to ai assistants

Chris Heilmann +2 · LIVE

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

Videos

See all

Related articles

See all