> Markdown version of [/jobs/ext/78083-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/78083-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Realm - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Bash Shell, Big Data, Computer Programming, Data Centers, DevOps, Distributed Systems, Python (Programming Language), Machine Learning, Open Source Technology, Reliability Engineering, Software Systems, Scripting, Computer Networking Systems, High Performance Computing, Computer Network Technologies, Deep Learning, Reliability of Systems, Kubernetes, Infrastructure Automation Frameworks, Bare Metal - **Published:** May 16, 2026 - **Apply:** https://find.jobs/jobs-near-me/senior-site-reliability-engineer-london/2771715732-2/ ## About the Role * Strong ownership mindset with focus on delivery and accountability * Experience building maintainable, well-documented systems in complex environments * Ability to operate effectively in ambiguous and rapidly evolving contexts * Clear and effective communication skills with collaborative, low-ego approach, * 5+ years of experience in site reliability engineering, DevOps, systems administration, or high-performance computing * Strong written and verbal communication skills in English * Experience deploying and operating container orchestration or workload scheduling systems (e.g. Kubernetes or similar) * Programming or scripting experience in Go, Python, or Bash * Familiarity with infrastructure automation and infrastructure-as-code tools * Strong technical foundation in computing or related discipline Preferred Experience * Experience operating large-scale machine learning or AI-compute workloads * Background in multi-tenant distributed systems at scale * Hands-on experience with data centre or bare-metal infrastructure * Knowledge of high-performance networking technologies * Experience managing large-scale storage systems (commercial or open-source) ## Description High-growth infrastructure company focused on delivering large-scale compute, data centre capacity, and power solutions for advanced machine learning workloads. Platforms support leading research and industry teams requiring high-performance computing at significant scale. Fast-paced environment with emphasis on ownership, execution speed, and quality. Culture centred on pragmatic problem-solving, cross-functional collaboration, and full lifecycle responsibility., * Position operating across software, infrastructure, and operations to ensure reliability, scalability, and performance of a globally distributed compute platform. * Close collaboration with networking, platform engineering, and physical infrastructure teams to design and operate systems supporting high-demand computational workloads. * Hands-on engineering role requiring strong systems expertise, with responsibility for resolving complex production issues, improving system resilience, and enhancing platform observability. Responsibilities * Deployment and management of large-scale compute clusters using automation tooling, with adaptation to customer requirements * Validation and optimisation of compute, storage, and networking systems in coordination with internal teams and vendors * Execution of large-scale data migrations between cloud and on-premise environments with focus on efficiency and cost * Troubleshooting across the full stack, including hardware, networking, and distributed systems * Development of internal tooling and automation to improve deployment speed, reliability, and operational efficiency Participation in an on-call rotation required (approximately one week per month). ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)