Site Reliability Engineer - Observability

Instaffo GmbH
Hamburg, Germany
17 days ago
Apply on de.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English, German
Job source

Tech stack

Artificial Intelligence Amazon Web Services Linux Open Source Technology Reliability Engineering Prometheus Grafana

Job description

Please note that this position requires work authorization for Germany. The language requirements for this position are: English - Fluent, German - Fluent.

We use AI to accelerate - not to replace thinking. We design the system, steer the output, and take responsibility for what we ship. Fast where it makes sense. Careful where it matters.

Activities

Take full ownership of smartclip’s internal utility and platform tooling. Focus your energy on the intersection of observability, automation, and developer infrastructure. Don’t just maintain existing systems - evolve them, research cutting-edge open-source alternatives, and implement them.

Forget expensive enterprise SaaS. Invest in deep in-house expertise. Understand our systems end-to-end, maintain total flexibility, and contribute back to the open-source ecosystem we depend on.

Face these challenges:

  • Build & Evolve: Operate and advance our observability stack (including Prometheus, Grafana, and Forgejo).
  • Go Open Source First: Replace “buy” decisions with robust “build & maintain” strategies.
  • Engineer the Platform: Design observability as a platform capability. Define SLOs and create actionable alerting to stop incidents before they start.
  • Secure the Stack: Embed security engineering into the delivery process. Find vulnerabilities before the pen tests do.
  • Master the Infrastructure: Navigate Linux systems and distributed tooling. Balance bold exploration with production stability., * Ownership over tickets: You’re trusted with real responsibility, not just tasks. No unnecessary bureaucracy, no micromanagement - we rely on you to take things forward.
  • Build > Talk: We test what works - not what sounds good. Fail fast, learn faster.
  • High standards, low ego: We take our work seriously, but not ourselves. Direct feedback, honest collaboration, no drama.
  • Stay sharp: Hackathons, conferences, community - we invest in your growth and keep you at the cutting edge.
  • And yes - the fundamentals are covered too: 30 days of vacation + Dec 24 & 31 off, Smart Fridays (4 days week possible), mobility (Germany ticket & JobRad), sports & health offerings, mental health support, corporate benefits, RTL+ access, and more.

Requirements

Be motivated by systems thinking and deep technical curiosity. Stop being a consumer - start being a builder., * Design and evolve production-grade setups on GCP or AWS.

  • Show us your contributions to open-source projects.
  • Turn your passion for root-cause analysis into blameless post-mortems.

About the company

About smartclip

smartclip is the adtech development unit of RTL Group - Europe’s leading free-to-air broadcaster group. Our proprietary advertising technology is custom-built for the needs of European broadcasters and publishers - enabling media owners to implement smarter monetisation strategies. We are committed to delivering the most innovative ad experiences spanning in-stream, out-stream, addressable TV, connected TV, audio, and gaming - ultimately empowering brands with true cross-screen storytelling opportunities on all devices.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · World Congress 2025

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

Videos

See all

Related articles

See all