Sovereign AI Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
- Design and run Kubernetes environments optimized for AI inference, retrieval, experimentation, and agent execution in secure or isolated settings
- Deploy and operate open-source or open-weight model stacks, model gateways, vector databases, and supporting platform components
- Build reproducible platform automation using Infrastructure as Code and GitOps approaches for stable, auditable delivery
- Manage local registries, package mirrors, secrets, access controls, storage, networking, and observability in environments with limited or no public cloud dependency.
- Optimize GPU, compute, and storage usage for reliable AI workloads while maintaining security and data sovereignty requirements
Examples of market tools, models, and platform components expected
- Inference and local serving stacks such as vLLM, Ollama, llama.cpp, or OpenAI-compatible self-hosted endpoints.
- Open-source or open-weight models appropriate for sovereign deployment, for example coding-capable and general-purpose families hosted internally through approved serving layers.
- Platform tooling such as Kubernetes, Helm, Terraform, Ansible, ArgoCD, private registries, Qdrant or similar vector stores, and Open WebUI or comparable internal interfaces.
- Developer-facing integration options such as VS Code-compatible extensions, Continue-style local model connectors, or editor integrations pointed at internal APIs instead of external SaaS endpoints.
- Hardware awareness covering GPU-backed nodes, CPU-only fallback options, storage performance, network isolation, and on-prem or dedicated infrastructure patterns.
Requirements
- 5+ years in platform engineering, DevOps, SRE, or MLOps, with strong Kubernetes and Linux expertise.
- Proven experience with AI infrastructure, model serving, private or on-prem deployments, and production operations for LLM-based workloads.
- Strong hands-on skills in Python plus automation tooling such as Terraform, Ansible, Helm, and GitOps workflows.
- Good understanding of networking, storage, access control, monitoring, and operational hardening in high-security environments.
- Comfortable working in sovereignty-driven environments where auditability, isolation, and controlled data handling are mandatory.
About the company
T-Systems is part of the Deutsche Telekom Group, with around 30.000 employees worldwide. We create technology with purpose to generate a positive impact on society. We are looking for curious talent, eager to learn, take on challenges, and contribute ideas that transform our customers’ experience.
We trust people: we offer autonomy, continuous support, and a collaborative environment where you can grow without limits. We are one global team, guided by respect, integrity, and a passion for doing better every day.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
MLOps – What’s the deal behind it?
Stephan Gillich - Bringing AI Everywhere
How to Become an AI Engineer