Applied AI Engineer, Site Reliability Engineer - EMEA

Mistral Ai
München, Germany
18 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cyber Security Computer Networks Software Debugging DevOps Python (Programming Language) Linux Kernel Systems Development Life Cycle Role-Based Access Control
+14 more
Reliability Engineering Ansible Prometheus Runbook Google Cloud Enterprise Software Applications Large Language Models Grafana Software Security Multi-Cloud Kubernetes Free and Open-Source Software Terraform Golang

Job description

The Applied AI team is Mistral’s customer-facing technical organization. We work directly with enterprise clients from pre-sales through implementation to deploy cutting-edge AI solutions that deliver measurable business impact.

Our team combines deep ML expertise with strong customer engagement skills, operating like startup CTOs who own end-to-end project execution. Our SRE team works transversally across customer engagements, enabling value creation through Mistral tech - at scale. By joining the team you will bridge the gap between cutting-edge AI research and real-world enterprise applications, ensuring our solutions are robust, scalable, and aligned with both customer needs and Mistral’s technological vision.

About The Job

You will be one of the founding engineers of the Applied AI SRE sub-team.

Your mission, alongside the team, is to build and operate the framework to ensure Mistral’s solution delivery is reliable and sustainable - and applied uniformly across all our accounts, both Mistral-hosted and customer-hosted. You should already have a strong understanding of what operational excellence looks like, and you’re ready to scale your impact.

You will operate in four concurrent modes:

  • BUILD - Design for a fleet of Mistral platforms and apps. Build proactivity to reduce reactivity. Productize reliability, author runbooks, create SLO templates, implement observability.

  • RUN - Operate the Tier-1 customer environments that Mistral are contracted to operate. Ensure SLO compliance, own on-call and incident response, manage drift, partner with Technical Support as L3 escalation, champion high signal post-mortems.

  • ENABLE - Productize how Mistral deploy, secure, and scale our Applied AI solutions. Engineer on-demand provisioning, author security baseline packages, embed security guardrails, automate everything.

  • SECURE - Own the security operations layer for our customer-side deployments. Lead CVE response across the fleet, ship supply-chain integrity controls (SBOM, signed images, provenance), co-page with InfoSec on security incidents, enforce secure-config baselines.

This is a framework-first, fleet management role at heart. If you’re excited by the difference between solving one customer’s problem and structurally solving the class of problem for every customer, this is the role.

How We Work in Applied AI

  • We care about people and outputs.
  • What matters is what you ship, not the time you spend on it
  • Bureaucracy is where urgency goes to vanish. You talk to whoever you need to talk to. The best idea wins, whether it comes from a principal engineer or someone in their first week.
  • Always ask why. The best solutions come from deep understanding, not from copying what worked before
  • We say what we mean. Feedback is direct, timely, and given because we care.
  • No politics. Low ego, high standards.
  • We embrace an unstructured environment and find joy in it.

Requirements

  • Fluent in English.
  • 5+ years in SRE, Production Engineering, or DevOps, with a record of shipping tooling.

  • Strong multi-tenant Kubernetes fluency, namespace segmentation, network policy, RBAC, admission control, operations at scale.

  • On-call discipline: incident response, blameless post-mortem culture, runbook-first mindset.

  • Observability stack in production: Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Signoz.

  • Infrastructure as code: Terraform, Ansible (or close equivalents).

  • Proficient in Python and/or Golang for tooling and automation.

  • Security mindset: you treat secure-SDLC, CVE response, and supply-chain integrity as reliability properties of the shipped artifact, not as someone else’s job.

  • Strong written communication skills: runbooks, post-mortems, and customer-facing incident comms are core deliverables of this role.

  • Comfortable operating with high autonomy in an ambiguous, fast-paced environment - and disciplined enough to defend the team’s scope when work tries to spill in.

  • Solid Linux internals, networking debug, and distributed-systems fundamentals., * Cloud or application security background (AppSec, K8s security, supply chain - SBOM, cosign, SLSA). At least one of our early hires must bring this; if it’s you, flag it.

  • Experience operating LLM / model-serving stacks in production

  • Experience with multi-cloud or on-prem hybrid customer environments (AWS, GCP, Azure, sovereign clouds).

  • Open-source contributions, particularly in SRE, observability, or security tooling.

Benefits & conditions

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

About the company

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems-across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector-co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

Videos

See all

Related articles

See all