> Markdown version of [/jobs/ext/97644-cloud-network-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/97644-cloud-network-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Network Reliability Engineer - **Company:** Apple Inc. - **Location:** Sunnyvale, CA, United States - **Salary:** $212,000.0 - $318,400.0 - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Systems Engineering, Cloud Computing, Distributed Systems, Fault Tolerance, Protocol Buffers, Monitoring of Systems, Hardware Virtualization, OSI Models, JSON, Network Configuration and Change Management, Network Architecture, Network Control, Network Service, OpenStack, Service Discovery, Software Engineering, System Programming, Extensible Markup Language (XML), Data Logging, Cloud-native Network Functions (CNF), Multithreading, Concurrency, Mttr, Multi-Cloud, Caching, Kubernetes, Infrastructure Automation Frameworks, SDN Network, Low Latency, Build Tools, Api Design, Restful APIs - **Published:** May 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=34291cca7e19084b ## About the Role Do you have experience in Systems engineering?, Expert knowledge of API design and interface technologies (JSON, ProtoBuf, REST, RPC, XML, etc) In depth knowledge of K8s, OpenStack, system virtualization, build systems and infrastructure as code Strong knowledge of observability systems (metrics, logging, tracing) and qualification engineering. Broad knowledge of networking solutions across OSI layers 3 through 7. Excellent written and verbal communication skills with the ability to clearly articulate risk, reliability trade-offs, and operational priorities. Proven ability to manage competing priorities, drive initiatives to completion, and deliver results in fast-paced environments. Minimum Qualifications Extensive experience in software engineering, systems engineering, or infrastructure engineering. Strong background in designing, operating, and supporting highly available, fault-tolerant distributed systems at hyper scale. Strong systems programming skills including multi-threading, concurrency, caching, batching Solid understanding of network infrastructure and software-defined networking (SDN). Ability to lead cross-functional collaboration and influence technical decisions across teams. ## Description We are seeking an experienced and visionary Cloud Network Reliability Engineer to drive the technical strategy and execution for ensuring the availability, performance, scalability, and resiliency of Apple's global network services. In this role, you will work as a technical leader solving complex networking challenges at massive scale, partnering with engineering, infrastructure, and operations teams across Apple to deliver reliable, fault-tolerant systems.., As a technical leader within the Cloud Networking organization, you will define and drive the reliability and resiliency architecture for Apple's network platform services. You will be responsible for establishing SRE and SWE best practices, architecting fault-tolerant network control and data planes, and championing data-driven decision-making through observability and automation. You will drive resilient cloud networking solutions that operate reliably across multiple cloud providers and global regions, handling failures gracefully and maintaining service availability. Your technical leadership will ensure Apple's network services meet demanding availability, latency, resilience, and security requirements while continuously improving operational maturity. We are looking for a technical expert who deeply understands cloud networking at scale, is passionate about operating mission-critical, globally distributed infrastructure, preventing outages through proactive engineering, and driving long-term reliability improvements through architectural excellence. ","responsibilities":"Define and drive the long-term technical vision, architecture, and reliability strategy for large-scale cloud networking platforms spanning control plane and data plane systems. Architect and evolve fault-tolerant, highly available network services, ensuring graceful degradation and consistent performance under partial and systemic failure scenarios. Establish platform-wide resiliency patterns including service discovery, health checking, automated failover, rate limiting, circuit breaking, and traffic management across multi-region and multi-cloud environments. Lead the design of network configuration management, routing state distribution, traffic engineering, and capacity planning systems, balancing scalability, correctness, and operational simplicity. Serve as a senior technical authority and architectural reviewer, influencing critical design decisions across multiple teams and ensuring network failure modes are explicitly addressed. Build and champion automation-first reliability solutions, including topology discovery, deployment safety mechanisms, self-healing systems, and operational tooling that reduce toil and improve MTTR. Define and own reliability metrics and observability standards (SLIs, SLOs, error budgets), using data to drive engineering trade-offs, reliability investments, and incident response improvements. Multiply impact through cross-team technical leadership, embedding reliability early in design, mentoring engineers, and sharing deep technical knowledge through documentation and technical talks. ## Related Videos - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Top Must-Visit Developer Conferences in the US in 2026](https://www.wearedevelopers.com/magazine/679-top-must-visit-developer-conferences-in-the-us-in-2026) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)