C005025 Kubernetes Platform Operations and Maintenance Expert (NS) - THU 20 Aug

EMW, Inc.
Vijfhoek (Brussel), Belgium
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Computer Networks Software Debugging Disaster Recovery Key Management Network Architecture Nginx Node.Js Role-Based Access Control Prometheus Runbook Software Engineering Zabbix
+6 more
Data Logging SUSE Linux Grafana Kubernetes Commvault Vulnerability Analysis

Job description

The Kubernetes Platform Operations and Maintenance Expert will report to the Head of Processing and Storage Section.

Deliver monthly artifacts; maintain cluster operations; perform upgrades and patching; maintain documentation; provide knowledge transfer; ensure compliance with defined performance SLA and internal procedures.

Participate in weekly status update meetings, activity planning and other meetings as instructed, physically in the office, or in person via electronic means using Conference Call capabilities, according to the Team Leaders instructions.

Tooling stack (inside the cluster): The Kubernetes Platform Operations and Maintenance Expert operate and maintain the following Kubernetes native tooling: Prometheus & Grafana; Logging stack; NGINX or ingress; Cilium CNI; Backup tooling; On prem registry (e.g. Harbor/SUSE); ArgoCD GitOps; Metrics server; Cert manager; Zabbix integration hooks.

Out of Scope: Underlying VMs, storage, and network infrastructure (NCIA responsibility); disaster Recovery tests (NCIA responsibility; Service Contractor provides performance artifacts); application development or debugging; any activity outside the Kubernetes clusters.

Deliverables:

  • Cluster Health Report - node health, pod health, workloads, resource utilization, cluster events, degraded components.
  • Security Compliance Report - CIS benchmark checks, RBAC audit, network policy compliance, vulnerability scan results, remediation actions.
  • Capacity and Performance Report - CPU/memory/storage trends, saturation risks, performance anomalies, capacity forecasting.
  • Patching and Upgrade Report - Applied patches, pending patches, quarterly upgrade status, upgrade readiness checks.
  • Incident and RCA Report - Summary of P1-P3 incidents, root cause analysis, corrective actions, prevention measures.
  • Configuration Drift Report - Comparison of cluster configuration vs approved baseline, drift items, remediation plan.
  • Monthly Change Log - All changes applied to clusters, manifests, GitOps commits, registry updates, configuration changes.
  • Monthly Risk Register Update - Identified risks, severity, mitigation actions, owner, expected resolution timeline.

Enabling activities:

  • Daily monitoring of cluster health and alerts.
  • Node lifecycle operations (join, drain, cordon, replacement).
  • Management of Kubernetes components (CNI, CSI, ingress, metrics server, CoreDNS).
  • Certificate rotation and secrets management.
  • RBAC administration and periodic audits.
  • Network policy management.
  • Vulnerability scanning and remediation inside the cluster.
  • Backup operations using Commvault Velero/K10 (excluding DR tests).
  • Maintaining documentation and runbooks.
  • Assisting application teams with deployment troubleshooting.
  • Participation in planning meetings, retrospectives, and workshops.

Requirements

Skills, Knowledge & Experience:

  • The candidate must have a currently active NATO SECRET security clearance
  • University Degree + 3 years relevant experience, or equivalent.
  • 3+ years hands on Kubernetes operations experience.
  • Experience with SUSE RKE2 Kubernetes distributions.
  • Experience with monitoring, logging, repositories, ingress, CNI, and backup tooling, * Prior experience of working in an international environment comprising both military and civilian elements;
  • Knowledge of NATO responsibilities and organization, including ACO and ACT.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.be

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

4:40 min

Assessing common Kubernetes security incidents and misconfigurations

Rico Komenda Rico Komenda · World Congress 2025

5:02 min

Manual port forwarding configuration using network address translation

Oliver Seitz Oliver Seitz · World Congress 2025

1:39 min

Introduction to Kubernetes security challenges and opportunities

Marc Nimmerrichter · World Congress 2022

Videos

See all

Related articles

See all