Head of Platform Reliability
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Head of Platform Reliability Chain infrastructure carrying regulated money, in production, for institutions that hold it to the highest possible standard. You would own its reliability. That means defining the service levels, setting the error budget policy, and holding the authority to stop a release when the budget is spent - a call that stands regardless of seniority. You would build and lead the engineering team that runs it, on a rotation you design. We would rather you designed it properly than inherited something and patched it. Chain platforms fail differently from the systems most reliability engineers have run. State is expensive, restarts are not a strategy, and a surprising amount of what looks like infrastructure turns out to be protocol behaviour. If that sounds interesting rather than daunting, we should talk.
Requirements
- To have led site reliability or production engineering somewhere failure had consequences.
- To have run error budgets in practice - used one to stop a release, and defended the call
- Honest experience of continuous coverage - what burns people out, what does not, and the difference between a rotation on paper and one people can live with
-
Kubernetes, Terraform, and an observability stack you have actually debugged, rather than configured. Useful, not essential
- Blockchain nodes in production.
- Banking, payments, or somewhere else an auditor asked you to prove something.
Benefits & conditions
$165,000-$165,000 per year
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship
Why Blockchain? A Developer’s Perspective
The Geometry of Incidents: Connecting User Impact to Architecture
Never delegate the understanding