> Markdown version of [/jobs/ext/402292-senior-ai-devops-llmops](https://www.wearedevelopers.com/jobs/ext/402292-senior-ai-devops-llmops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI DevOps / LLMOps - **Company:** TechBiz Global GmbH - **Location:** Darmstadt, Germany (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Engineering, Computer Clusters, Nvidia CUDA, Continuous Integration, DevOps, Github, IBM Hardware Management Console, Monitoring of Systems, Ansible, Management of Software Versions, Virtualization Technology, AI Infrastructure, Policy as Code, Pulumi, Delivery Pipeline, Large Language Models, Multi-Agent Systems, HybridCloud, Gitlab-ci, Machine Learning Operations, Terraform - **Published:** June 21, 2026 - **Apply:** https://de.indeed.com/viewjob?jk=4933ad249afc35f4 ## About the Role Do you have experience in Virtualization?, Do you have a Master's degree?, * CI/CD & IaC: Expertise in GitHub Actions/GitLab CI, and Terraform or Pulumi. * AI Tooling: Experience with Weights & Biases, MLflow, LangSmith, or Arize Phoenix. * Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises hardware management. * Security: Familiarity with Open Policy Agent (OPA) and secret management (Vault). Experience: * 10+ years in DevOps, SRE, or Cloud Engineering. * 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs from notebook to production. * Proven experience managing Hybrid Cloud environments (e.g., AWS/Azure + Private ## Description * Automation of Build-to-Production * Design and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code. * Develop specialized workflows for PromptOps, ensuring that system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code. * Automate the deployment of Agentic workflows, managing the complexities of stateful AI interactions and multi-agent handoffs. 2. AI Infrastructure as Code (IaC) * Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, Pulumi, or Ansible. * Define and enforce Policy-as-Code for AI endpoints to ensure compliance with security, cost-usage limits, and data residency requirements. * Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless parity between On-Premises development and Cloud production. 3. Safe Experimentation & Controlled Releases * Architect Progressive Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing (where new models run in parallel with production to compare outputs). * Build "Evaluation-in-the-Loop" gates within the pipeline to automatically test for bias, hallucination, and performance degradation before a release. * Implement A/B testing frameworks specifically designed for LLM outputs and agentic behavior. 4. Monitoring & Observability - Establish deep observability into Inference Endpoints, tracking metrics like tokens-per- second, latency, and drift in model accuracy. * Integrate feedback loops that capture production "edge cases" to feed back into the ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)