Site Reliability Engineer (SRE)

Deutsche Telekom AG
Düsseldorf, Germany
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
German

Tech stack

Cloud Engineering Cyber Security Data Retention Linux Network Connections Routing Reliability Engineering Ansible Prometheus YAML Grafana Infrastructure as Code (IaC)
+6 more
Containerization Kubernetes Infrastructure Automation Frameworks Deployment Automation Data Management Terraform

Job description

Experteer Overview You will design, implement, and operate a high-performance observability platform for distributed Edge and Fog infrastructures. You’ll build dashboards, configure alerts, and optimize monitoring stacks to ensure stability, availability, and efficiency in resource-constrained environments. The role emphasizes cloud-native tools, IaC, and cross-team collaboration to support digital transformation and reliable operations. This position offers a modern, collaborative context within a mission-critical tech environment that powers resilient edge computing. Pay / Benefits * Design, implement, and maintain a high-performance monitoring stack based on Prometheus, Grafana, Loki, and Alertmanager for distributed Fog and Edge infrastructures * Develop specialized dashboards for hardware health, resource utilization, and network connectivity * Design and implement alerting strategies for autonomous operating models and air-gapped environments * Implement local data retention concepts, log rotation, and efficient data management for resource-constrained environments * Optimize monitoring stack performance considering limited CPU, memory, and storage on Fog nodes * Integrate Kubernetes monitoring solutions and analyze operational metrics to identify optimization opportunities * Apply Infrastructure as Code (IaC) and YAML configurations to automate deployment and management of monitoring components * Contribute to evolving a robust, scalable observability platform for Edge and Fog computing environments Tasks * Strong experience with Prometheus (PromQL, Federation, Remote Write, Local Retention, dashboards) * Advanced Grafana skills (dashboards, alerting, provisioning as code) * Experience with Loki (LogQL, retention, compaction, alerting/incident management) * Solid knowledge of Alertmanager (routing, inhibition, escalation) * Kubernetes & cloud-native technologies experience; operating Kubernetes environments * Understanding of container platforms and cloud-native architectures * Edge & Fog computing experience in resource-constrained environments; offline/air-gapped scenarios * YAML proficiency; Infrastructure as Code exposure * Strong IT security principles and problem-solving approach * Excellent German language skills (written and spoken, C1) * Experience in Site Reliability Engineering or infrastructure operations * Nice-to-have: IaC tools (Terraform, Ansible); GitOps familiarity; Linux admin; OpenTelemetry; mission-critical infra Key requirements * flexible working hours * hybrid work model * competitive salary * development opportunities * modern work environment * health and well-being benefits

Requirements

maintain log rotation, and efficient data management for resource-constrained environments * Optimize monitoring stack performance considering limited CPU, memory, and storage on Fog nodes * Integrate Kubernetes monitoring solutions and analyze operational metrics to identify optimization opportunities * Apply Infrastructure as Code (IaC) and YAML configurations to automate deployment and management of monitoring components * Contribute to evolving a robust, scalable observability platform for Edge and Fog computing environments Tasks * Strong experience with Prometheus (PromQL, Federation, Remote Write, Local Retention, dashboards) * Advanced Grafana skills (dashboards, alerting, provisioning as code) * Experience with Loki (LogQL, retention, compaction, alerting/incident management) * Solid knowledge of Alertmanager (routing, inhibition, escalation) * Kubernetes & cloud-native technologies experience; operating Kubernetes environments * Understanding of container platforms and aaaaaaaa Code architectures * Edge & Fog computing experience in resource-constrained environments; offline/air-gapped scenarios * YAML proficiency; Infrastructure as Code exposure * Strong IT security principles and problem-solving approach * Excellent German language skills (written and spoken, C1) * Experience in Site Reliability Engineering or infrastructure operations * Nice-to-have: IaC tools (Terraform, Ansible); GitOps familiarity; Linux admin; OpenTelemetry; mission-critical infra Key requirements * flexible working hours * hybrid work model * competitive salary * development opportunities * modern work environment * health and well-being benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eu.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:47 min

Exploring career opportunities and recruitment open positions

Kurt Eder · LIVE

2:00 min

Introduction to YAML syntax and basic formatting

Chris Ayers · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

Videos

See all

Related articles

See all