Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
Experteer Overview You will design, implement, and operate a high-performance observability platform for distributed Edge and Fog infrastructures. You’ll build dashboards, configure alerts, and optimize monitoring stacks to ensure stability, availability, and efficiency in resource-constrained environments. The role emphasizes cloud-native tools, IaC, and cross-team collaboration to support digital transformation and reliable operations. This position offers a modern, collaborative context within a mission-critical tech environment that powers resilient edge computing. Pay / Benefits * Design, implement, and maintain a high-performance monitoring stack based on Prometheus, Grafana, Loki, and Alertmanager for distributed Fog and Edge infrastructures * Develop specialized dashboards for hardware health, resource utilization, and network connectivity * Design and implement alerting strategies for autonomous operating models and air-gapped environments * Implement local data retention concepts, log rotation, and efficient data management for resource-constrained environments * Optimize monitoring stack performance considering limited CPU, memory, and storage on Fog nodes * Integrate Kubernetes monitoring solutions and analyze operational metrics to identify optimization opportunities * Apply Infrastructure as Code (IaC) and YAML configurations to automate deployment and management of monitoring components * Contribute to evolving a robust, scalable observability platform for Edge and Fog computing environments Tasks * Strong experience with Prometheus (PromQL, Federation, Remote Write, Local Retention, dashboards) * Advanced Grafana skills (dashboards, alerting, provisioning as code) * Experience with Loki (LogQL, retention, compaction, alerting/incident management) * Solid knowledge of Alertmanager (routing, inhibition, escalation) * Kubernetes & cloud-native technologies experience; operating Kubernetes environments * Understanding of container platforms and cloud-native architectures * Edge & Fog computing experience in resource-constrained environments; offline/air-gapped scenarios * YAML proficiency; Infrastructure as Code exposure * Strong IT security principles and problem-solving approach * Excellent German language skills (written and spoken, C1) * Experience in Site Reliability Engineering or infrastructure operations * Nice-to-have: IaC tools (Terraform, Ansible); GitOps familiarity; Linux admin; OpenTelemetry; mission-critical infra Key requirements * flexible working hours * hybrid work model * competitive salary * development opportunities * modern work environment * health and well-being benefits
Requirements
maintain log rotation, and efficient data management for resource-constrained environments * Optimize monitoring stack performance considering limited CPU, memory, and storage on Fog nodes * Integrate Kubernetes monitoring solutions and analyze operational metrics to identify optimization opportunities * Apply Infrastructure as Code (IaC) and YAML configurations to automate deployment and management of monitoring components * Contribute to evolving a robust, scalable observability platform for Edge and Fog computing environments Tasks * Strong experience with Prometheus (PromQL, Federation, Remote Write, Local Retention, dashboards) * Advanced Grafana skills (dashboards, alerting, provisioning as code) * Experience with Loki (LogQL, retention, compaction, alerting/incident management) * Solid knowledge of Alertmanager (routing, inhibition, escalation) * Kubernetes & cloud-native technologies experience; operating Kubernetes environments * Understanding of container platforms and aaaaaaaa Code architectures * Edge & Fog computing experience in resource-constrained environments; offline/air-gapped scenarios * YAML proficiency; Infrastructure as Code exposure * Strong IT security principles and problem-solving approach * Excellent German language skills (written and spoken, C1) * Experience in Site Reliability Engineering or infrastructure operations * Nice-to-have: IaC tools (Terraform, Ansible); GitOps familiarity; Linux admin; OpenTelemetry; mission-critical infra Key requirements * flexible working hours * hybrid work model * competitive salary * development opportunities * modern work environment * health and well-being benefits
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on eu.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Finding Jobs in Germany
Where to Find German Tech Jobs
Fully Remote Software Engineer Jobs
How to Find Tech Jobs in Berlin