Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
We are seeking a Site Reliability Engineer (SRE) with a strong background in observability, secure logging, and automation. The ideal candidate will have hands-on experience with Elasticsearch and/or Prometheus platforms. This role encompasses critical responsibilities in platform operations, including incident management, execution of scheduled maintenance, and contributing to engineering tasks focused on enhancing system stability. The SRE will also be responsible for adhering to standard operating procedures (SOPs) and actively contributing to their continuous improvement by providing constructive feedback., * Platform Engineering & DevOps: Manage Kubernetes and container orchestration, including Helm chart configurations and CI/CD pipelines (Jenkins, ArgoCD). Develop automation scripts (Python, Bash, Go) and deploy Infrastructure-as-Code (IaC) solutions.
- Observability, Monitoring & Visualisation: Maintain Prometheus solutions (scrape configurations, alert rules, PromQL queries), administer Thanos and Grafana.
- Elastic Stack Operations & Log Management: Configure and optimise Elasticsearch clusters, Logstash pipelines, and Kibana dashboards for secure, scalable log processing.
- Incident Response, Troubleshooting & Collaboration: Participate in 24x7 on-call rotations for rapid incident response, troubleshoot platform, data and performance issues, and engage in Major Incident Management (MIM).
- Secure Operations & Compliance: Ensure system operations meet security and data protection requirements, maintain secure documentation, and manage access control policies.
Requirements
- Strong grasp of Linux concepts, preferably in Kubernetes environments.
- Solid understanding of networking fundamentals and REST APIs.
- Proficiency in Python, Go, or Bash.
- Proficiency in Git-based configuration management workflows.
- Strong experience with the ELK Stack, Prometheus, and Grafana.
- Experience with Elasticsearch and/or OpenSearch.
- Knowledge of Kubernetes and familiarity with cloud platforms.
- Familiarity with CI/CD tools like Helm, Jenkins, or ArgoCD.
- Good understanding of PromQL is an advantage.
- Fluent English communication skills.
- Willingness to work shift-based 24x7 on-call support, including weekends and holidays.
- Must possess Γ2 security clearance.
- Citizenship required: Member state of the EU and NATO. No dual citizenship outside these countries.
- Must reside in Germany and hold a German labor contract.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleβ¦