Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments. This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating things rather than repeatedly fixing them manually. The role sits at the intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective. What youâll be doing
Youâll build Python-based automation around incident management, operational runbooks and routine infrastructure tasks. Youâll integrate monitoring, infrastructure and ITSM platforms through APIs, helping improve the quality of alerts through better correlation, enrichment, suppression and deduplication. Youâll also develop internal tools, command-line utilities, dashboards and potentially ChatOps capabilities that allow operational teams to resolve issues more quickly. A major part of the role will be taking existing operational processes and asking: âWhy are we still doing this manually?â Youâll then design a safe, controlled and auditable way of automating it., ServiceNow, Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation. This is not an AI/ML development position. Weâre looking for someone who understands production infrastructure and can use software engineering and automation to make that infrastructure more reliable. Youâll be joining a growing organisation where youâll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment.
Requirements
You should have good commercial experience in Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure operations, together with strong hands-on Python automation skills. Youâll also need experience with:
- Monitoring and observability tools such as Prometheus, Grafana or similar
- Production incident management and/or on-call environments
- Automating operational runbooks and repetitive infrastructure processes
- APIs and systems integration
- Version-controlled automation and operational tooling
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where To Find Software Engineering Jobs
Why Upskilling And Reskilling is Important For Developers
Dev Digest 120 - Apple and peers
Fully Remote Software Engineer Jobs