> Markdown version of [/jobs/ext/2729759-platform-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2729759-platform-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Site Reliability Engineer - **Company:** Radiant - **Location:** UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Automation of Tests, Bash Shell, Ubuntu (Operating System), System Configuration, Dynamic Host Configuration Protocol, Linux, File Systems, Domain Name System (DNS), Python (Programming Language), Linux System Administration, Routing, Performance Tuning, Ansible, Prometheus, Software Deployment, TCP/IP, Virtual Local Area Networks, AI Infrastructure, Computer Networking Systems, Grafana, Kubernetes, Information Technology, Bare Metal - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/platform-site-reliability-engineer-radiant-8352629 ## About the Role * 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model in an SRE or equivalent role * 3+ years experience in both running, deploying and optimising orchestration platforms with a strong emphasis on Kubernetes * Expert-level Linux administration, especially Ubuntu distributions * Proficiency in system tuning, disk I/O optimization, and hardware-level performance tweaks * Strong networking fundamentals: TCP/IP, DNS, DHCP, VLANs, routing, switching * Strong experience with API interrogation * Strong experience with infrastructure scripting and automation (Bash, Python, Ansible) * Deep understanding of observability principles and tools (Prometheus, Grafana preferred) * Strong grasp of ITSM and service operation best practices * Excellent communication and mentorship skills * Comfortable interfacing with internal stakeholders and external customers * Bonus: Knowledge of running AI workloads via orchestration platforms Bonus Requirements * Bachelor or Masters Level degree in Computer Science, Engineering or related field, or equivalent experience. * LPIC Certifications * ITIL Foundation level qualification or equivalent experience * Certified Kubernetes Administrator (CKA) Qualities we look for: * You approach problems with a systems mindset - balancing practical execution with long-term scalability * You elevate the team, setting high standards for technical quality and engineering excellence. * You hold yourself and others accountable - giving direct feedback and expecting the same * You take initiative, owning challenges end-to-end and proactively driving solutions. * You invest in others, mentoring to build both capability and confidence. ## Description We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure powers GPU-native workloads, multi-tenant control planes, and high-performance AI systems designed for the most demanding environments. We are not building a generic cloud. We are building purpose-built AI infrastructure - from powered land, to compute, to software . As we scale our platform and expand our engineering organisation, we are looking for leaders who can build strong teams, uphold high standards, and deliver reliably at pace. Role Responsibilities * Deploy and Manage Kubernetes Clusters, deployed at scale to support AI centric workloads, across both our bare metal clusters and via trusted partner infrastructure * Develop Kubernetes Manifests and Operators: Facilitate application deployments and maintain Kubernetes-native services for networking, storage, security, identity and infrastructure management * Optimize Linux system configuration including kernel, driver, filesystem and services to support workloads running via our orchestration layer * Build and maintain automation scripts and infrastructure as code to support platform lifecycle, as well as simplifying troubleshooting for Incident resolution and provision of tooling for our support organisation * Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. * Maintain and enhance Radiant's observability stack: Prometheus, Grafana, and custom monitoring integrations * Operate and support services in 24x7 production environments, including on-call rotation * Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations * Mentor junior engineers and act as an Operational requirements consultant to other departments * Communicate technical decisions clearly to non-technical stakeholders and customers * Uphold a culture of: do, document, automate * Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks. * Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)