> Markdown version of [/jobs/ext/3083954-ai-infrastructure-platform-operations-engineer-remote-in-the-eu](https://www.wearedevelopers.com/jobs/ext/3083954-ai-infrastructure-platform-operations-engineer-remote-in-the-eu). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure & Platform Operations Engineer (remote in the EU) - **Company:** Mirantis - **Location:** Berlin, Germany - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Cloud Computing, Computer Networks, Distributed Systems, InfiniBand, Linux System Administration, Reliability Engineering, Prometheus, Runbook, AI Infrastructure, Graphics Processing Unit (GPU), Computer Network Operations, High Performance Computing, Grafana, Kubernetes, Infrastructure Automation Frameworks, Hardware Infrastructure - **Published:** September 26, 2026 - **Apply:** https://jobs.smartrecruiters.com/Mirantis/744000139217287-ai-infrastructure-platform-operations-engineer-remote-in-the-eu- ## About the Role * 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles. * Strong Linux administration and troubleshooting skills. * Good understanding of networking concepts and experience diagnosing infrastructure-related issues. * Working knowledge of Kubernetes in production environments. * Experience supporting production infrastructure and services. * Strong analytical and problem-solving skills. * Experience working within structured operational and incident management processes. * Excellent communication and collaboration skills. * Ability to work within a shift-based operational environment. Experience in one or more of the following areas is highly desirable: * NVIDIA GPU infrastructure and accelerated computing platforms. * InfiniBand networking and NVIDIA UFM. * Kubernetes platform operations. * AI infrastructure or HPC environments. * Site Reliability Engineering (SRE) or Platform Engineering. * Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry. * Infrastructure automation technologies and Infrastructure-as-Code practices. * Large-scale distributed systems and production platforms. ## Description We are building a European AI Infrastructure & Platform Operations team responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs, high-performance networking, Kubernetes, and next-generation platform technologies. The team is responsible for ensuring the availability, performance, and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. Working at the intersection of infrastructure, networking, and platform operations, you will help support the environments that power modern AI workloads. This is an opportunity to work with some of the latest technologies in AI infrastructure while contributing to the evolution of AI-powered operational services through platforms such as k0rdent AI., * Monitor, operate, and support production AI infrastructure platforms. * Investigate and resolve infrastructure, networking, hardware, and platform-related incidents. * Support NVIDIA GPU infrastructure and associated platform services. * Monitor and troubleshoot Kubernetes-based environments. * Investigate performance, availability, and reliability issues across infrastructure and platform components. * Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues. * Participate in incident response, root cause analysis, and operational improvement activities. * Contribute to improvements in monitoring, observability, automation, and operational processes. * Maintain operational documentation, runbooks, and knowledge articles. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [From foundation model to hosted AI solution in minutes](https://www.wearedevelopers.com/videos/1170-from-foundation-model-to-hosted-ai-solution-in-minutes) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)