(Senior) DevOps Engineer

Blockbrain GmbH
Hamburg, Germany
11 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Shift work
Languages
English, German
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Software as a Service Continuous Integration DevOps Github Python (Programming Language) Key Management Reliability Engineering Prometheus TypeScript
+10 more
Software Vulnerability Management Policy as Code Scripting Istio Large Language Models Grafana Kubernetes Information Technology Terraform Dynatrace

Job description

You join the Platform team and own the reliability of the systems our AI agents run on. This is a reliability engineering role: you treat operations as a software problem, so where others run a manual procedure, you write the automation that makes it unnecessary. You define what “healthy” means in numbers, measure it, and hold the line on it in production. When something breaks, you bring the system back and then make sure it cannot break the same way twice.

Reliability & Operations

  • Define and own SLOs and SLIs for the platform and manage error budgets against them
  • Carry on-call, act as incident commander, and run blameless post-incident reviews that produce real follow-up
  • Run production readiness and capacity planning ahead of demand, not after the page fires

Platform & Infrastructure

  • Run and harden the Kubernetes platform (Helm, GitOps, service mesh) and the cloud underneath it (Terraform, multi-region)
  • Own observability: metrics, logs, and distributed tracing, so problems surface before users feel them

Automation & Efficiency

  • Eliminate toil through automation, self-healing systems, and automated remediation
  • Drive cost visibility and FinOps practice across cloud and LLM spend

Security

  • Bake security into the platform: least-privilege access, secrets management, policy as code, and vulnerability management, * Flexible Work Models: Full-time position on-site in Stuttgart, Hamburg or Munich (3 days per week) with flexible working hours.
  • Benefits: Deutschland-Ticket, Wellpass fitness membership, access to the latest AI tools, and regular company and team off-sites.
  • Top Team: International team with exceptional talents.
  • High Growth Potential: Steep learning curve in a fast-growing AI startup. A high level of personal responsibility and the freedom to actively shape processes.
  • Top Equipment: MacBook, iPhone, headset, and all the tools you need to perform at your best.

Requirements

  • Professional Experience: 5+ years in DevOps, SRE, or platform engineering, ideally operating production SaaS at scale. Hands-on experience running Kubernetes in production is essential.
  • Communication: Clear and calm under pressure. Can coordinate an incident and write a post-mortem others learn from. Works in English; German is a plus.
  • Tech Affinity: Treats infrastructure as code and operations as a software discipline. Genuinely enjoys automating manual work away.
  • Solution Orientation: Measures success in incidents that did not happen. Fixes root causes, not symptoms.
  • Organizational Talent: Plans capacity and reliability work ahead of demand and balances on-call, project work, and toil reduction.
  • Education: Degree in computer science or a related field, or equivalent hands-on experience. We care about what you can operate, not the certificate.
  • Hard Skills: Kubernetes, Terraform / IaC, CI/CD (e.g. GitHub Actions), observability (Prometheus, Grafana, distributed tracing), a major cloud (AWS, Azure, or GCP), scripting (Python, Go, or TypeScript), secrets management, and policy as code.
  • Soft Skills & Mindset: Strong ownership, blameless culture, a bias toward automation, calm in incidents, and a security-by-default mindset.

About the company

At Blockbrain, we don’t just talk about AI - we use it every day. In this role, you will:

  • Use coding agents to build automation, write infrastructure code, and reason through failure modes faster
  • Automate operational toil and incident workflows with AI in the loop
  • Collaborate with the product team to give real-world feedback on Blockbrain’s own tools from an operator’s perspective
  • Stay curious about emerging AI capabilities and apply them to platform and reliability work

We’re not looking for AI experts - we’re looking for people who are genuinely open to working with AI as a daily co-pilot.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch ¡ World Congress 2026 Europe

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar ¡ World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters ¡ World Congress 2023

1:17 min

Storyblok engineering context and tech stack overview

Sebastian Gierlinger Sebastian Gierlinger ¡ World Congress 2025

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur ¡ LIVE

7:15 min

Installing Istio programmatically with bash scripts

Thomas SßdbrÜcker ¡ LIVE

Videos

See all

Related articles

See all