Sovereign Cloud Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+24 more
Job description
You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement.
| The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment., Kubernetes | Gardener | Helm | Jenkins | ArgoCD | Git | IaC | Python | Go | Bash | Prometheus | PromQL | Thanos | OpenTelemetry | Grafana | Elasticsearch | OpenSearch | Logstash | Kibana | ServiceNow | PagerDuty | REST APIs |
Kubernetes & Platform Operations
· Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components.
· Troubleshoot availability, performance and deployment issues.
· Ensure secure, scalable, resilient and highly available platforms.
CI/CD & Automation
· Operate and optimize Jenkins and ArgoCD pipelines for automated deployments.
· Implement IaC and Git-based deployment workflows.
· Develop automation and operational tools using Python, Go and/or Bash.
· Automate provisioning, health/compliance checks, alerting and reporting.
Monitoring & Observability
· Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL.
· Develop and maintain Grafana dashboards and continuously improve monitoring and alerting.
Logging & Log Management
· Operate and optimize Elasticsearch/OpenSearch, Logstash and Kibana.
· Monitor log ingestion, storage, performance and reliability.
· Support centralized troubleshooting, anomaly detection, security and compliance.
Integration & Operations
· Integrate observability and logging platforms with ServiceNow, PagerDuty and other enterprise tools via secure APIs.
· Handle operational requests and participate in Scrum, DevOps and service-improvement activities.
· Collaborate with internal teams, SAP, suppliers and stakeholders.
Incident & Problem Management
· Participate in a 24/7 on-call and shift rotation, including weekends and public holidays.
· Respond to platform, monitoring, logging and deployment incidents.
· Perform RCA, support Major Incident Management (MIM) and implement sustainable corrective actions.
· Continuously improve platform stability, resilience and operational processes.
Security, Compliance & Documentation
· Operate platforms in accordance with EU data protection requirements, GDPR, VS-NfD, VSA/GHB, and relevant enterprise security standards.
· Enforce access-control policies and regularly review permissions.
· Maintain secure, version-controlled technical documentation.
· Maintain architecture diagrams, configuration documentation, operational procedures, and runbooks.
· Develop and maintain knowledge-management materials.
Requirements
The candidate must hold valid citizenship in a country that is a full member of both the European Union (EU) and NATO., Must hold a valid and verifiable Ü2 security clearance in accordance with the German Security Clearance Act (Sicherheitsüberprüfungsgesetz - SÜG) and applicable preventive personnel sabotage-protection requirements., · Several years of professional experience as an SRE, DevOps Engineer, Platform Engineer, or similar.
· Strong Linux knowledge, particularly in Kubernetes-based environments.
· Hands-on experience with Kubernetes, preferably Gardener.
· Strong experience with Helm and Kubernetes configuration management.
· Practical experience with Jenkins and ArgoCD.
· Strong understanding of Infrastructure as Code (IaC) principles.
· Programming/scripting skills in Python, Go, and/or Bash.
· Strong understanding of Git-based configuration and deployment workflows.
· Solid knowledge of networking fundamentals and REST APIs.
· Experience with Prometheus, PromQL, Grafana, Thanos, and/or OpenTelemetry.
· Hands-on experience with Elasticsearch/OpenSearch, Logstash, and Kibana.
· Experience integrating enterprise monitoring and incident-management platforms via APIs.
· Strong analytical and problem-solving skills.
· Ability to work effectively in a highly operational and international environment.
· Fluent English communication skills, written and spoken.
· German language skills are advantageous.
Preferred Certifications
· Elastic Certified Engineer
· LPIC-2
· Certified Kubernetes Administrator (CKA)
Operational & Security Requirements
· Willingness to participate in a 24x7 on-call/shift rotation, including weekends and public holidays.
· Willingness to actively participate in incident response and Major Incident Management.
Benefits & conditions
If the candidate holds multiple citizenships, all citizenships must be from countries that are members of both the EU and NATO.
Employment in Germany
The candidate must:
· Be employed directly by a legal entity registered in Germany.
· Hold a German employment contract.
· Be subject exclusively to German labor law.
· Comply with all applicable German tax, social-security, and employment regulations.
· Reside in Germany.
Employment through non-German entities, including foreign subcontractors or affiliates, does not meet this requirement.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Backend Developer Salary in Germany [2023]
Fullstack developer salary in Germany [2023]
Jobs in Germany for Americans
Frontend Developer Salary in Germany [2023]