Site Reliability Engineer - ClickHouse
Role details
Job location
Tech stack
Job description
We run one of the largest self-managed ClickHouse installations on AWS, at petabyte scale, and we're actively preparing it for the next 10-50× of growth. This role sits at the centre of that effort. You won't be in a typical "keep the lights on" SRE role. The work is about turning a fast-growing, stateful system into a predictable, well-automated platform (provisioning, scaling, rebalancing, recovery). That means reducing operational stress, designing safe automation for data-heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort. You'll work on the kind of problems that only show up at large scale (petabytes of data, thousands of cores, constant ingestion).
- Managing large fleets of EC2-based VMs, disks, and networking for data-intensive workloads.
- Improving operational tooling around deploys, schema changes, backups, restores, and incident response.
- Working closely with ClickHouse engineers to turn database-level needs into infra-level solutions.
- Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation.
- Participating in on-call and incident response, with a strong focus on making incidents rarer over time.
- Designing and automating, not just responding to alerts.
You should join this team if you like deep ownership of production systems, and are not afraid of working with stateful infrastructure.
Requirements
We're looking for people (EU/UK based) that like deep ownership of production systems, people that are not afraid of working with stateful infrastructure and love working in AWS, VMs, automation, and making messy systems reliable., * Prior experience with ClickHouse or other OLAP databases.
- Strong experience operating production infrastructure on AWS.
- Hands-on experience with VM-based systems (EC2), not just managed PaaS.
- Experience automating infrastructure using tools like Terraform, Ansible, or similar.
- Solid understanding of Linux systems (disk, memory, networking, failure modes).
- Experience supporting stateful systems (databases, queues, storage systems, etc.).
- Ability to debug and reason about performance and reliability issues in production.
- You're comfortable owning systems end-to-end, including on-call responsibilities.
You don't need to be a ClickHouse expert on day one. We'll teach you the database internals, but you do need to enjoy owning complex infrastructure.
Benefits & conditions
- Enthusiastic drivers. We need proactive people that can fully own projects and get them done, and know to get help when needed. "Are we there yet?" is the wrong question.
- Optimistic problem solvers. Things get hard here sometimes, whether it's scaling, shipping complex products, handling a stream of support requests, or trying to ship something that touches multiple teams. We need people who won't get disheartened, and will collaborate, iterate, and ship their way out of anything.
- Grown ups. We're an international bunch of weirdos, but one thing unites us: everyone is kind, considerate, and professional towards each other. This isn't about age or experience, it's about being low-ego, flexible, and respectful.
- Genuine builders. PostHog is full of people who just love building stuff, people who would still be building software even if there wasn't a paycheck at the end. If this sounds like you, we should talk.