Sovereign Cloud Engineer

TMT Prüfservice GmbH & Co. KG
Walldorf, Germany
17 days ago
Apply on www.adzuna.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Languages
English, German
Job source

Tech stack

Application Programming Interfaces (APIs) Bash Shell Cloud Computing Cloud Engineering Computer Programming Continuous Integration DevOps Elasticsearch Monitoring of Systems Python (Programming Language) Scrum Methodology Reliability Engineering
+22 more
Logstash Prometheus SAP (Applications) Systems Integration Data Logging Scripting Cloud Platform System Data Ingestion System Availability Delivery Pipeline Grafana Kubernetes Helm Charts Infrastructure as Code (IaC) Git Kubernetes Infrastructure Automation Frameworks Deployment Automation Kibana Restful APIs Pagerduty Jenkins Servicenow

Job description

For our growing team, we are looking for an experienced Sovereign Cloud / Site Reliability Engineer (SRE) to support the secure, reliable, and continuous operation of modern cloud-native and Kubernetes platforms.

You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement.

The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment., * Be employed directly by a legal entity registered in Germany.

  • Hold a German employment contract.
  • Be subject exclusively to German labor law.
  • Comply with all applicable German tax, social-security, and employment regulations.
  • Reside in Germany.

**Employment through non-German entities, including foreign subcontractors or affiliates, does not meet this requirement. **

Security Clearance - Ü2

You must hold a valid and verifiable Ü2 security clearance in accordance with the German Security Clearance Act (Sicherheitsüberprüfungsgesetz - SÜG) and applicable preventive personnel sabotage-protection requirements.

Tasks

Kubernetes & Platform Operations

  • Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components.
  • Troubleshoot availability, performance and deployment issues.
  • Ensure secure, scalable, resilient and highly available platforms.

CI/CD & Automation

  • Operate and optimize Jenkins and ArgoCD pipelines for automated deployments.
  • Implement Infrastructure as Code (IaC) and Git-based deployment workflows.
  • Develop automation and operational tools using Python, Go and/or Bash.
  • Automate provisioning, health and compliance checks, alerting and reporting.

Monitoring & Observability

  • Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL.
  • Develop and maintain Grafana dashboards.
  • Continuously improve monitoring and alerting capabilities.

Logging & Log Management

  • Operate and optimize Elasticsearch/OpenSearch, Logstash and Kibana.
  • Monitor log ingestion, storage, performance and reliability.
  • Support centralized troubleshooting, anomaly detection, security and compliance.

Integration & Operations

  • Integrate observability and logging platforms with ServiceNow, PagerDuty and other enterprise tools via secure APIs.
  • Handle operational requests.
  • Participate in Scrum, DevOps and service-improvement activities.
  • Collaborate with internal teams, SAP, suppliers and stakeholders.

Incident & Problem Management

  • Participate in a 24/7 on-call and shift rotation, including weekends and public holidays.
  • Respond to platform, monitoring, logging and deployment incidents.
  • Perform Root Cause Analysis (RCA).
  • Support Major Incident Management (MIM).
  • Implement sustainable corrective actions.
  • Continuously improve platform stability, resilience and operational processes.

Requirements

You must hold valid citizenship in a country that is a full member of both the European Union (EU) and NATO., Key Technologies

Kubernetes Gardener Helm Jenkins ArgoCD Git IaC Python Go Bash Prometheus PromQL Thanos OpenTelemetry Grafana Elasticsearch OpenSearch Logstash Kibana ServiceNow PagerDuty REST APIs

Technical Skills & Experience

  • Proven experience as a Sovereign Cloud Engineer and/or Site Reliability Engineer (SRE).
  • Strong experience operating and maintaining Kubernetes clusters, ideally with Gardener.
  • Hands-on experience with Kubernetes workloads, deployments, Helm charts and platform components.
  • Experience troubleshooting availability, performance and deployment issues.
  • Experience with Jenkins and ArgoCD for automated CI/CD deployments.
  • Solid understanding of Infrastructure as Code (IaC) and Git-based deployment workflows.
  • Programming or scripting experience with Python, Go and/or Bash.
  • Experience with automation of provisioning, health checks, compliance checks, alerting and reporting.
  • Hands-on experience with Prometheus, PromQL, Thanos and OpenTelemetry.
  • Experience developing and maintaining Grafana dashboards and monitoring/alerting solutions.
  • Experience with Elasticsearch/OpenSearch, Logstash and Kibana.
  • Understanding of log ingestion, storage, performance and reliability.
  • Experience supporting centralized troubleshooting, anomaly detection, security and compliance.
  • Experience integrating platforms with ServiceNow, PagerDuty and other enterprise tools via secure REST APIs.
  • Experience with incident management, Root Cause Analysis (RCA) and Major Incident Management (MIM).
  • Experience working in 24/7 operational environments, including on-call and shift rotations.
  • Strong understanding of platform security, scalability, resilience, reliability and high availability.
  • Experience working in Scrum and DevOps environments.
  • Strong collaboration skills and experience working with internal teams, SAP, suppliers and other stakeholders.

Languages

  • English: Required
  • German: A plus

Benefits & conditions

What You Can Expect

  • A technically challenging role within a highly secure and business-critical cloud environment.
  • The opportunity to work with modern cloud-native, Kubernetes and observability technologies.
  • An international working environment with teams and stakeholders across different locations.
  • The opportunity to contribute to automation, platform reliability and continuous improvement.
  • Remote / hybrid working possibilities within Germany.

We Look Forward to Hearing from You

Are you ready to bring your expertise to a challenging cloud engineering environment?

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · World Congress 2022

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:47 min

Exploring career opportunities and recruitment open positions

Kurt Eder · LIVE

3:27 min

Defining DevOps through its historical origins and foundational texts

Sonal Patil · LIVE

4:51 min

Executing simple full-text search queries using the Kibana interface

Derek Binkley · LIVE

Videos

See all

Related articles

See all