> Markdown version of [/jobs/ext/2011397-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2011397-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** Deutsche Telekom AG - **Location:** Düsseldorf, Germany - **Contract:** Permanent contract - **Skills:** Cloud Engineering, Cyber Security, Data Retention, Linux, Network Connections, Routing, Reliability Engineering, Ansible, Prometheus, YAML, Grafana, Infrastructure as Code (IaC), Containerization, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Data Management, Terraform - **Published:** August 10, 2026 - **Apply:** https://eu.experteer.com/career/view-jobs/site-reliability-engineer-sre-m-f-d-duesseldorf-nordrheinwestfalen-deutschland-58868636 ## About the Role maintain log rotation, and efficient data management for resource-constrained environments * Optimize monitoring stack performance considering limited CPU, memory, and storage on Fog nodes * Integrate Kubernetes monitoring solutions and analyze operational metrics to identify optimization opportunities * Apply Infrastructure as Code (IaC) and YAML configurations to automate deployment and management of monitoring components * Contribute to evolving a robust, scalable observability platform for Edge and Fog computing environments Tasks * Strong experience with Prometheus (PromQL, Federation, Remote Write, Local Retention, dashboards) * Advanced Grafana skills (dashboards, alerting, provisioning as code) * Experience with Loki (LogQL, retention, compaction, alerting/incident management) * Solid knowledge of Alertmanager (routing, inhibition, escalation) * Kubernetes & cloud-native technologies experience; operating Kubernetes environments * Understanding of container platforms and aaaaaaaa Code architectures * Edge & Fog computing experience in resource-constrained environments; offline/air-gapped scenarios * YAML proficiency; Infrastructure as Code exposure * Strong IT security principles and problem-solving approach * Excellent German language skills (written and spoken, C1) * Experience in Site Reliability Engineering or infrastructure operations * Nice-to-have: IaC tools (Terraform, Ansible); GitOps familiarity; Linux admin; OpenTelemetry; mission-critical infra Key requirements * flexible working hours * hybrid work model * competitive salary * development opportunities * modern work environment * health and well-being benefits ## Description Experteer Overview You will design, implement, and operate a high-performance observability platform for distributed Edge and Fog infrastructures. You'll build dashboards, configure alerts, and optimize monitoring stacks to ensure stability, availability, and efficiency in resource-constrained environments. The role emphasizes cloud-native tools, IaC, and cross-team collaboration to support digital transformation and reliable operations. This position offers a modern, collaborative context within a mission-critical tech environment that powers resilient edge computing. Pay / Benefits * Design, implement, and maintain a high-performance monitoring stack based on Prometheus, Grafana, Loki, and Alertmanager for distributed Fog and Edge infrastructures * Develop specialized dashboards for hardware health, resource utilization, and network connectivity * Design and implement alerting strategies for autonomous operating models and air-gapped environments * Implement local data retention concepts, log rotation, and efficient data management for resource-constrained environments * Optimize monitoring stack performance considering limited CPU, memory, and storage on Fog nodes * Integrate Kubernetes monitoring solutions and analyze operational metrics to identify optimization opportunities * Apply Infrastructure as Code (IaC) and YAML configurations to automate deployment and management of monitoring components * Contribute to evolving a robust, scalable observability platform for Edge and Fog computing environments Tasks * Strong experience with Prometheus (PromQL, Federation, Remote Write, Local Retention, dashboards) * Advanced Grafana skills (dashboards, alerting, provisioning as code) * Experience with Loki (LogQL, retention, compaction, alerting/incident management) * Solid knowledge of Alertmanager (routing, inhibition, escalation) * Kubernetes & cloud-native technologies experience; operating Kubernetes environments * Understanding of container platforms and cloud-native architectures * Edge & Fog computing experience in resource-constrained environments; offline/air-gapped scenarios * YAML proficiency; Infrastructure as Code exposure * Strong IT security principles and problem-solving approach * Excellent German language skills (written and spoken, C1) * Experience in Site Reliability Engineering or infrastructure operations * Nice-to-have: IaC tools (Terraform, Ansible); GitOps familiarity; Linux admin; OpenTelemetry; mission-critical infra Key requirements * flexible working hours * hybrid work model * competitive salary * development opportunities * modern work environment * health and well-being benefits ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Unlocking the potential of Digital & IT at Vodafone](https://www.wearedevelopers.com/videos/602-unlocking-the-potential-of-digital-it-at-vodafone) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [Finding Jobs in Germany](https://www.wearedevelopers.com/magazine/375-finding-jobs-in-germany) - [Where to Find German Tech Jobs](https://www.wearedevelopers.com/magazine/366-where-to-find-german-tech-jobs) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How to Find Tech Jobs in Berlin](https://www.wearedevelopers.com/magazine/291-how-to-find-tech-jobs-in-berlin) - [Finding IT & Technology English-speaking Jobs in Germany ](https://www.wearedevelopers.com/magazine/446-finding-it-technology-english-speaking-jobs-in-germany)