> Markdown version of [/jobs/ext/3587985-senior-platform-engineer-ai-observability](https://www.wearedevelopers.com/jobs/ext/3587985-senior-platform-engineer-ai-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform Engineer AI & Observability - **Company:** Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften - **Location:** Garching bei München, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Computer Engineering, Continuous Integration, Linux, DevOps, Python (Programming Language), Machine Learning, Prometheus, Security Information and Event Management, Data Logging, Cloud Platform System, High Performance Computing, Retrieval-Augmented Generation, Delivery Pipeline, Large Language Models, Grafana, Agentic-AI, Containerization, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Free and Open-Source Software, Ollama, Machine Learning Operations, VLLM, Model Inference, llama.cpp, Docker - **Published:** October 5, 2026 - **Apply:** https://www.adzuna.de/details/5910335095 ## About the Role Required/Minimum Qualifications * Master's degree (or equivalent) in Computer Science, Data Science, Computer Engineering, or a related field. Other Requirements * Experience with Linux, Docker, Kubernetes, and cloud-native technologies. * Knowledge of observability, monitoring, logging, tracing, and security concepts. * Programming and scripting skills, preferably in Python, Go, or Bash. * Experience with DevOps, MLOps, platform engineering, or infrastructure automation. * Strong communication, collaboration, and problem-solving skills. * Excellent written and spoken English. Additional or Preferred Qualifications: * Experience with AI, machine learning, LLMs, or AI-assisted operations. * Hands-on experience with inference frameworks such as vLLM, Triton, Ollama, or Llama.cpp. * Experience operating GPU-accelerated, large-scale Kubernetes, or HPC environments. * Knowledge of Prometheus, Grafana, Helm, GitOps, and CI/CD. * Familiarity with RAG architectures, vector databases, MCP-based services, or AI agents. * Experience with security monitoring tools (such as Falco or Tracee). * Experience in research projects or open-source software development. ## Description * Design, deploy, and operate scalable AI, observability, and cloud-native platforms based on Kubernetes and HPC technologies. * Build and optimize AI services, LLM inference platforms, and GPU-enabled workloads. * Develop and maintain monitoring, logging, tracing, and security solutions using open-source technologies. * Create standardized deployment workflows, automation, and platform best practices. * Enable reliable, secure, and multi-tenant operation of federated research infrastructures. * Collaborate with project partners and provide technical leadership in architecture, implementation, and operations.