> Markdown version of [/jobs/ext/3529904-senior-ai-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3529904-senior-ai-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # (Senior) AI Platform Engineer - **Company:** Evoila Gmbh - **Location:** Mainz, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Encodings, Continuous Integration, Data Security, Python (Programming Language), Octopus Deploy, Performance Tuning, Role-Based Access Control, Regression Testing, OpenAI, Ansible, Azure Machine Learning, Data Streaming, Load Balancing, Retrieval-Augmented Generation, Large Language Models, Agentic-AI, Kubernetes, Infrastructure Automation Frameworks, Machine Learning Operations, Nim (Programming Language), Artificial Intelligence Governance, Terraform, KServe, Golang - **Published:** October 2, 2026 - **Apply:** https://www.adzuna.de/details/5904003556 ## About the Role * Mindestens 3 Jahre Erfahrung im Betrieb und der Optimierung von Plattformen und Services at Scale in Kubernetes-Umgebungen, egal ob On-Premise oder in der Cloud. * Fundierte Kenntnisse in der Architektur und dem Betrieb von skalierbaren Lösungen unter Verwendung von Kubernetes, idealerweise inklusive GPU-Workloads. * Programmierkenntnisse in Python, idealerweise ergänzt um Golang oder Java. * Erfahrung mit mindestens einem Model-Serving-Framework (z. B. vLLM, NVIDIA Triton, KServe) sowie erste oder vertiefte Erfahrung mit MLOps-Werkzeugen (z. B. MLflow, Kubeflow, ClearML). * Erfahrung mit Infrastructure-as-Code-Tools wie Ansible und Terraform sowie CI/CD- und GitOps-Werkzeugen (z. B. Argo CD, Flux). * Sehr gute Kommunikationsfähigkeiten in Deutsch und Englisch. Idealerweise hast du bereits Kunden beraten und kannst technische Themen zielgruppengerecht vermitteln - vom Engineering-Team bis zum Management., * Mindestens 5 Jahre einschlägige Erfahrung im Plattform-Engineering. * Architektur-Ownership: Du verantwortest das Design von AI-Plattformen end-to-end und triffst technologische Grundsatzentscheidungen. * Technische Führung in Kundenprojekten (Lead-Rolle) sowie Mentoring von Kolleginnen und Kollegen. Darüber hinaus verfügst du idealerweise über: * Erfahrung mit Cloud-AI-Diensten (z. B. AWS SageMaker/Bedrock, Azure ML/Azure OpenAI/Azure AI Foundry) und deren Integration mit Kubernetes-Workloads. * Kenntnisse im Betrieb von LLM-basierten Architekturen, z. B. RAG-Pipelines, Vektordatenbanken und Embedding-Services sowie Agentic-AI-Ansätze und Model Context Protocol (MCP). * Verständnis von System-Level-Grundlagen des LLM-Servings (Rate Limiting, Token Streaming, Load Balancing) und LLM-Konzepten wie Reasoning, Tool Calling und Prompt Templates. * Erfahrung mit neuen Serving-Bausteinen wie llm-d oder Kubernetes Inference Gateway. * Kenntnisse in der Absicherung von AI-Plattformen, z. B. TLS, RBAC und Network Policies innerhalb von Kubernetes-Umgebungen, sowie Grundverständnis von AI-Governance (z. B. EU AI Act). * Zertifizierungen im Bereich relevanter Technologien (z. B. Certified Kubernetes Administrator/Developer, NVIDIA-Zertifizierungen). * Bereitschaft zu gelegentlichen Reisen im DACH-Raum. ## Description * Planung, Aufbau und Betrieb von hochverfügbaren AI/ML-Plattformen und -Services, insbesondere auf Basis von Kubernetes in On-Premises-, Hybrid- und Cloud-Umgebungen. * Design und Umsetzung skalierbarer GPU-Infrastrukturen innerhalb von Kubernetes-Clustern - inklusive GPU-Scheduling und -Sharing (z. B. Kueue, KAI-Scheduler, Run:ai, NVIDIA MIG/Time-Slicing) mit Blick auf Auslastung und Kosteneffizienz. * Aufbau und Betrieb von Model-Serving- und Inferenz-Diensten (z. B. vLLM, NVIDIA Triton, KServe, NVIDIA NIM, Ray) für klassische ML-Modelle und Large Language Models. * Entwicklung und Betrieb von MLOps-/LLMOps-Workflows für produktive GenAI-Services - inklusive Guardrails & Policies, Evaluierung, Regressionstests sowie Kosten- und Qualitätsmonitoring. * Bereitstellung von Fine-Tuning- und Re-Training-Workflows (z. B. LoRA, kontinuierliches Re-Training) auf der Plattform. * Monitoring, Troubleshooting und Performance-Tuning von AI-Plattformen (u. a. GPU-Auslastung, Inferenz-Latenz, Durchsatz), um höchste Performance und Verfügbarkeit sicherzustellen. * Beratung unserer Kunden in Bezug auf Best Practices, AI-Governance, Datensicherheit und Datenschutz im Kontext von AI-Plattformen auf Kubernetes. * Enge Zusammenarbeit mit Data-Science-, Data-Platform- und Entwicklungsteams zur Sicherstellung der nahtlosen Integration von AI-Lösungen. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [From AI Assistance to Agentic Systems: Scaling Sovereign AI in Banking](https://www.wearedevelopers.com/videos/100070-from-ai-assistance-to-agentic-systems-scaling-sovereign-ai-in-banking) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Develop AI-powered Applications with OpenAI Embeddings and Azure Search](https://www.wearedevelopers.com/videos/828-develop-ai-powered-applications-with-openai-embeddings-and-azure-search) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)