> Markdown version of [/jobs/ext/1126599-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1126599-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** gridscale GmbH - **Location:** Köln, Germany (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Border Gateway Protocol, Cursor (Graphical User Interface Elements), Software Debugging, Linux, DevOps, Firmware, Python (Programming Language), Kernel-Based Virtual Machine, Linux System Administration, Node.Js, OpenStack, Reliability Engineering, Markdown, Ansible, Prometheus, Virtual Local Area Networks, Datadog, Large Language Models, Grafana, Prompt Engineering, Git, Kubernetes, Bare Metal, Terraform, Automation Anywhere - **Published:** July 2, 2026 - **Apply:** https://de.indeed.com/viewjob?jk=8fb7c538bed3f847 ## About the Role OpenStack · Kubernetes · KVM · Linux · Bare-metal · Ansible · Terraform · FluxCD/ ArgoCD · Git · Go · Python * Claude Code/ Cursor/ agentic coding tooling, * Several years of hands-on experience running production infrastructure (SRE, Platform, or DevOps). * Solid OpenStack experience - deployed, operated, and debugged it in production. * End-to-end compute infrastructure management, from bare-metal lifecycle through hypervisor and virtual compute node operations (migration, host evacuation, graceful drains, capacity rebalancing). The skill matters more than the specific tooling - what counts is having done it at scale and automated it. * Strong with Infrastructure as Code (Ansible, Terraform) and GitOps (FluxCD or ArgoCD), plus solid Linux administration including on bare-metal. * Active, daily practice of AI-assisted engineering, with opinions formed from real use. You can describe a workflow where an LLM saved you half a day, and one where you should have skipped it. Theoretical interest doesn't count. * Fluent English, written and spoken - our team is distributed, and this is the working language. * Nice to Have + Production experience with Kubernetes and the cloud-native ecosystem. + Production-quality Go and/or Python. + Deeper agentic tooling craft (Claude Code, Cursor, Aider): custom agent / sub-agent setups, hooks, prompt engineering, your own workflows or skills and managing a Markdown-first knowledge base as substrate for AI workflows. + Advanced compute-node tuning (CPU pinning, NUMA, hugepages, SR-IOV / PCI passthrough) and basic network debugging (VLANs, BGP). + Observability tooling (Prometheus, Loki, Grafana, etc.) and auto-remediation / self-healing systems (StackStorm, Event-Driven Ansible, or similar). + Experience in security-critical environments and with edge or multi-site deployments. * Soft Skills + A continuous-improvement mindset and ownership for what you build. + You see AI tooling as a structural shift in how engineering gets done - not a trend, not a threat and want to shape how the team adopts it. + You enjoy sharing knowledge, learning from peers, and can synthesize ideas clearly. ## Description * Develop Infrastructure as Code with Ansible and Terraform - typically spec-first with LLM assistance, then human-validated; push this further via custom agent / sub-agent setups, agentic test generation, and prompt-engineered review loops. * Drive the ongoing development of our Kubernetes stack and GitOps workflows (FluxCD / ArgoCD). * Own the full lifecycle of our compute infrastructure - from bare-metal (firmware, provisioning, hardware health) through hypervisors to virtual compute nodes - and build the automation that keeps capacity healthy and rolls out updates without disturbing tenant workloads. * Build and extend the AI substrate that compounds our output: Markdown knowledge bases as retrieval substrate, agentic prototypes for incident triage and capacity planning, and deeper integration of agentic coding tools into daily work. * Contribute to the self-healing direction, turning today's manual runbooks into tomorrow's reasoning agents. Auto-remediation isn't a separate team here - it's how platform work is meant to land. * Design and implement test suites aligned with functional and technical specs (non-regression, performance, security). * Document and package the solution so users can deploy and operate it without friction, and keep improving the platform based on telemetry and user feedback. * Act as a technical reference and mentor across automation, platform engineering, and AI-tooling topics. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 119 - ❤️ === ❤️](https://www.wearedevelopers.com/magazine/454-dev-digest-119) - [Dev Digest 108 - Git off my cloud!](https://www.wearedevelopers.com/magazine/407-dev-digest-108-git-off-my-cloud)