Site Reliability Engineer (On-prem Kubernetes Operations) for NATO with security clearance

Wlg
Brussel, Belgium
14 days ago
Apply on www.careerjet.be
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Data Analysis Computer Networks DevOps Key Management Network Architecture Nginx Role-Based Access Control Reliability Engineering Prometheus Software Engineering Zabbix Data Logging
+5 more
SUSE Linux Grafana Kubernetes CIS Benchmarks Vulnerability Analysis

Job description

WLG and our selected partners use cookies and similar technologies (together “cookies”) that are necessary to present this website, and to ensure you get the best experience of it. If you consent to it, we will also use cookies for analytics and marketing purposes. See our to read more about the cookies we set. You can withdraw and manage your consent at any time, by clicking “Manage cookies” at the bottom of each website page. common–cookies–alert#acceptAll”>Accept all cookies common–cookies–alert#disableAll”>Decline all non-necessary cookies common–cookies–alert#openPreferences”>Cookie preferences common–cookies–preferences#open” class=”company-links bg-company-primary text-company-primary shadow-form inset-x-4 mb-4 rounded max-w-screen-sm p-4 overflow-auto max-h-[80vh] z-cookie-alert backdrop:bg-[rgba(0,0,0,0.6)] sm:p-10 sm:mb-auto fixed top-2”> Select which cookies you accept On this site, we always set cookies that are strictly necessary, meaning they are necessary for the site to function properly. If you consent to it, we will also set other types of cookies. You can provide or withdraw your consent to the different types of cookies using the toggles below. You can change or withdraw your consent at any time, by clicking the link “Manage Cookies”, which is always available at the bottom of the site. To learn more about what the different types of cookies do, how your data is used when they are set etc, see our . Strictly necessary These cookies are necessary to make the site work properly, and are always set when you visit the site. Vendors Teamtailor common–cookies–preferences#handleCategoryToggle”> Analytics These cookies collect information to help us understand how the site is being used. Vendors Teamtailor common–cookies–preferences#handleCategoryToggle”> Marketing These cookies are used to make advertising messages more relevant to you. In some cases, they also deliver additional functions on the site. Vendors Youtube common–cookies–preferences#handleAccept”>Accept these cookies common–cookies–preferences#handleDecline”>Decline all non-necessary cookies common–header–menu#toggle” data-common–header–menu-target=”button” > Career menu Employee Candidate Homepage common–dropdown#toggle”> Share page common–share#handleClick” data-provider=”Facebook”> Facebook common–share#handleClick” data-provider=”Twitter”> X common–share#handleClick” data-provider=”LinkedIn”> LinkedIn DevOps · Brussels Site Reliability Engineer (On-prem Kubernetes Operations) for NATO with security clearance Run the on-premise Kubernetes clusters behind a NATO organisation’s services in Brussels - operations, upgrades, security and the reporting that proves it. careersite–jobs–form-overlay#showFormOverlay” data-careersite–jobs–form-overlay-target=”coverButton”> Apply for this job blocks–cover–scroll#handleScrollDown” title=”Scroll to content”> A NATO organisation in Brussels runs its own Kubernetes platform on premise, and it is looking for the person who keeps it healthy. This is a pure platform-operations role: the clusters, the tooling inside them, and the monthly evidence that the whole thing is compliant and under control. Application development and the underlying virtual, storage and network infrastructure sit with other teams. What you would be doing

  • Daily monitoring of cluster health and alerts, and node lifecycle operations - join, drain, cordon, replacement.
  • Managing the Kubernetes components: CNI, CSI, ingress, metrics server and CoreDNS.
  • Certificate rotation and secrets management, RBAC administration with periodic audits, and network policy management.
  • Upgrades and patching, including quarterly upgrade readiness.
  • Vulnerability scanning and remediation inside the cluster.
  • Backup operations with Velero or K10 (disaster recovery testing stays with the customer).
  • Maintaining documentation and runbooks, and helping application teams troubleshoot their deployments.
  • Producing the monthly artefacts: cluster health, security compliance (CIS benchmarks, RBAC audit, network policy, vulnerabilities), capacity and performance, patching and upgrades, incidents and root cause analysis, configuration drift against the approved baseline, the change log and the risk register.

The stack inside the cluster

  • Prometheus and Grafana, a logging stack, NGINX or another ingress, Cilium CNI.
  • Backup tooling, an on-premise registry such as Harbor or SUSE, ArgoCD for GitOps, metrics server, cert-manager and Zabbix integration hooks.

Requirements

  • A university degree plus three years of relevant experience, or the equivalent.
  • Three or more years of hands-on Kubernetes operations.
  • Experience with SUSE RKE2.
  • Real experience of monitoring, logging, registries, ingress, CNI and backup tooling.
  • Nice to have: an international environment with both military and civilian colleagues, and an understanding of NATO’s structure.

Practical detail

  • Brussels, Belgium - on site; most daily activity has to be done in the office. No travel.
  • 28 September to 31 December 2026, 570 hours in total.
  • You must hold a valid clearance and be a citizen of a NATO member country with the right to work in Belgium.

If you would rather operate a platform properly than build a new one every quarter, this is the kind of role where that is the whole point.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.be
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt ¡ LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 ¡ LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum ¡ World Congress 2026 Europe

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz ¡ World Congress 2026 Europe

3:37 min

Why differing legacy workflows complicate monitoring tool migrations

Mathias Palmersheim Mathias Palmersheim ¡ Europe 2026 Virtual

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark ¡ LIVE

Videos

See all

Related articles

See all