Site Reliability Engineer - Observability & Internal Tools
Role details
Job location
Tech stack
Job description
We don't just subscribe to services; we engineer platforms. We believe that deep in-house expertise and a strong commitment to open source provide the flexibility and performance that enterprise SaaS cannot match. We design, we steer, and we own our stack end-to-end., You will be the guardian and architect of smartclip's internal infrastructure. Your goal is to turn observability and automation from "tools we use" into a "platform capability" that empowers every other engineer in the company.
- Own the Observability Stack: Take the lead on Prometheus, Grafana, and Forgejo. You won't just maintain them; you will evolve them into a world-class monitoring ecosystem.
- Engineer for Reliability: Design actionable alerting and define SLOs that move us from reactive firefighting to proactive stability.
- Champion Open Source: Evaluate and implement cutting-edge open-source alternatives to proprietary software. You decide what enters our stack and how it integrates.
- Secure the Pipeline: Integrate security engineering directly into our delivery process. You find the vulnerabilities before the pen tests do.
- Master the Metal: Navigate the depths of Linux systems and distributed tooling, balancing bold experimentation with rock-solid production stability., * Ownership over tickets: You're trusted with real responsibility, not just tasks. No unnecessary bureaucracy, no micromanagement - we rely on you to take things forward.
- Build > Talk: We test what works - not what sounds good. Fail fast, learn faster.
- High standards, low ego: We take our work seriously, but not ourselves. Direct feedback, honest collaboration, no drama.
- Stay sharp: Hackathons, conferences, community - we invest in your growth and keep you at the cutting edge.
- And yes - the fundamentals are covered too: 30 days of vacation + Dec 24 & 31 off, Smart Fridays (4 days week possible), mobility (Germany ticket & JobRad), sports & health offerings, mental health support, corporate benefits, RTL+ access, and more.
Requirements
We are looking for a "builder" who is bored by simple configuration and thrives on systems thinking., * Observability Expertise: You have a proven track record of implementing metrics, logs, and traces. You know how to turn noisy data into actionable insights.
- Linux & Systems Engineering: You are comfortable in the terminal and understand how distributed systems communicate and fail.
- Automation Mindset: You hate doing the same thing twice. You use code to automate infrastructure and eliminate toil.
- Ownership Culture: You embrace the "you build it, you run it" philosophy and take pride in the stability of the systems you touch.
The Power-Ups (Nice-to-haves):
- Hands-on experience with GCP or AWS at a production scale.
- Active contributions to the open-source community or a portfolio of self-hosted projects.
- Experience in conducting blameless post-mortems and driving root-cause analysis.