Site Reliability Engineer - ClickHouse

ClickHouse
Saint Pölten, Austria
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Saint Pölten, Austria

Tech stack

Amazon Web Services (AWS)
Amazon Web Services (AWS)
Databases
Software Debugging
Linux
Online Analytical Processing
Platform as a Service (PAAS)
Reliability Engineering
Ansible
Vertica
Terraform

Job description

We run one of the largest self-managed ClickHouse installations on AWS, at petabyte scale, and we're actively preparing it for the next 10-50× of growth. This role sits at the centre of that effort. You won't be in a typical "keep the lights on" SRE role. The work is about turning a fast-growing, stateful system into a predictable, well-automated platform (provisioning, scaling, rebalancing, recovery). That means reducing operational stress, designing safe automation for data-heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort. You'll work on the kind of problems that only show up at large scale (petabytes of data, thousands of cores, constant ingestion).

  • Managing large fleets of EC2-based VMs, disks, and networking for data-intensive workloads.
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response.
  • Working closely with ClickHouse engineers to turn database-level needs into infra-level solutions.
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation.
  • Participating in on-call and incident response, with a strong focus on making incidents rarer over time.
  • Designing and automating, not just responding to alerts.

You should join this team if you like deep ownership of production systems, and are not afraid of working with stateful infrastructure.

Requirements

We're looking for people (EU/UK based) that like deep ownership of production systems, people that are not afraid of working with stateful infrastructure and love working in AWS, VMs, automation, and making messy systems reliable., * Prior experience with ClickHouse or other OLAP databases.

  • Strong experience operating production infrastructure on AWS.
  • Hands-on experience with VM-based systems (EC2), not just managed PaaS.
  • Experience automating infrastructure using tools like Terraform, Ansible, or similar.
  • Solid understanding of Linux systems (disk, memory, networking, failure modes).
  • Experience supporting stateful systems (databases, queues, storage systems, etc.).
  • Ability to debug and reason about performance and reliability issues in production.
  • You're comfortable owning systems end-to-end, including on-call responsibilities.

You don't need to be a ClickHouse expert on day one. We'll teach you the database internals, but you do need to enjoy owning complex infrastructure.

Benefits & conditions

  • Enthusiastic drivers. We need proactive people that can fully own projects and get them done, and know to get help when needed. "Are we there yet?" is the wrong question.
  • Optimistic problem solvers. Things get hard here sometimes, whether it's scaling, shipping complex products, handling a stream of support requests, or trying to ship something that touches multiple teams. We need people who won't get disheartened, and will collaborate, iterate, and ship their way out of anything.
  • Grown ups. We're an international bunch of weirdos, but one thing unites us: everyone is kind, considerate, and professional towards each other. This isn't about age or experience, it's about being low-ego, flexible, and respectful.
  • Genuine builders. PostHog is full of people who just love building stuff, people who would still be building software even if there wasn't a paycheck at the end. If this sounds like you, we should talk.

Apply for this position