> Markdown version of [/jobs/ext/1579741-site-reliability-engineer-icloud](https://www.wearedevelopers.com/jobs/ext/1579741-site-reliability-engineer-icloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, iCloud - **Company:** Apple Inc. - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Apple Xcode, Linux, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Internet Control Message Protocol, Subnetting, OSI Models, Python (Programming Language), Network Protocols, Nginx, Open Source Technology, Systems Development Life Cycle, Reliability Engineering, Prometheus, Software Engineering, TCP/IP, Load Balancing, Grafana, Reliability of Systems, HybridCloud, Kubernetes, Information Technology, Splunk, Dynatrace, Docker - **Published:** July 16, 2026 - **Apply:** https://www.totaljobs.com/job/site-reliability-engineer/apple-job107696782 ## About the Role * Strong sense of ownership, customer service, and integrity proven through clear communication. * BS in Computer Science or related field, or equivalent employment * 5 + years experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment * Strong experience with deploying, supporting and supervising new and existing services, platforms, and application stacks * Experience with scale testing, disaster recovery, and capacity planning * Experience with observability platforms with Splunk, Grafana, Prometheus. * Demonstrable fluency in at least one of the following languages: Java, Python, or Go. * Experience with Kubernetes, Nginx, Envoy, Prometheus, and/or Docker., * Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model, Subnetting and Load Balancing strategies. * Understanding of the Linux Operating System, including Kernel, Memory, Process, Threads, Static / Shared Libraries, IPC, Signals. * Experience in developing iOS apps using Xcode and Swift. * Experience in OpenTelemetry Standards / distributed tracing like jaeger ## Description Apple Services' scale is BIG. Operating at our scale, across multiple geographies and servicing hundreds of millions of users presents unique challenges. As a Software Developer in SRE at Apple, you'll need to solve these problems using data, teamwork, and your own expertise. ASE Products Site Reliability teams are responsible for the reliability and performance of the server software stack that powers products like iCloud Photos, Mail, Drive, Backup and many more. We do that by focusing on reliability best practices from service inception to production, collaborating deeply with product development teams to deliver a superlative product and shared vision while leveraging data and automation as first principles. We run a mix of open source, vendor licensed, and internally developed tools to manage the end to end SDLC of our products. You'll learn these tools and have opportunities to improve them., * Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions. * Operate, monitor, and triage all aspects of our production and non-production environments. * Collaborate on code, infrastructure, design reviews, and process enhancements * Evaluate and integrate new technologies to improve system reliability, security, and performance. * Develop and implement automation to provision, configure, deploy, and monitor Apple services. * Participate in an oncall rotation providing hands-on technical expertise during service impacting events. * Contribute to capacity planning, scale testing, and disaster recovery exercises * Approach operational problems with a software engineering mindset. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Post-Quantum Cryptography: Preparing for Q-Day](https://www.wearedevelopers.com/videos/100179-post-quantum-cryptography-preparing-for-q-day) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 119 - ❤️ === ❤️](https://www.wearedevelopers.com/magazine/454-dev-digest-119)