> Markdown version of [/jobs/ext/2383057-software-engineer-infrastructure-services](https://www.wearedevelopers.com/jobs/ext/2383057-software-engineer-infrastructure-services). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Infrastructure Services - **Company:** Apple Inc. - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Cloud Computing, Code Review, Computer Networks, Computer Engineering, Software Design Documents, Software Design Patterns, Distributed Systems, Fault Tolerance, Python (Programming Language), Routing, Network Service, Reliability Engineering, Software Engineering, Concurrency, Reliability of Systems, Backend, Information Technology, SDN Network - **Published:** August 28, 2026 - **Apply:** https://www.dice.com/job-detail/3c941289-1737-4429-9166-0c796ec0e6be ## About the Role Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience 4 to 6+ years of professional software engineering experience designing, building, and operating production-grade distributed systems and backend infrastructure Deep foundation in computer science fundamentals, including distributed systems architecture, concurrency models, graph algorithms, and systems design Strong proficiency in at least one systems-level or high-performance language, such as Go, C++, Rust, or Python Direct experience designing and building fault-tolerant mechanisms, including automated self-healing, active remediation, circuit breaking, load shedding, and blast-radius mitigation for network services Demonstrated ability to model complex failure domains, handle network partitions and split-brain scenarios, and reason rigorously about system behavior under extreme load and degradation Practical experience with chaos engineering, fault injection, simulation-based testing, and stress testing in production or staging environments Track record of technical ownership, including authoring design documents, driving code reviews, and leading post-mortem root cause analyses Preferred Qualifications Master's or Ph.D. in Computer Science, Distributed Systems, Networking, or a related technical discipline Deep domain knowledge in cloud networking architectures, Software-Defined Networking (SDN) control planes, L3/L4 routing protocols, overlay networks, and Linux networking constructs (such as eBPF, XDP, or OVS) Hands-on experience implementing or tuning distributed consensus algorithms (such as Raft or Paxos) and closed-loop control systems or reconciliation controllers Experience architecting high-throughput telemetry pipelines and automated anomaly detection systems for real-time network health analysis Proven ability to digest academic papers and industry research, translating state-of-the-art resilience concepts into production systems History of notable open-source contributions, technical publications, or patent filings in distributed networking and systems reliability ## Description This is not a traditional network engineering role. This is a software engineering role for builders who love deep technical challenges, have strong fundamentals in distributed systems and fault-tolerant design, and want to turn cutting-edge ideas into production systems running at massive scale. We welcome early-career engineers with exceptional technical depth, if you have the intellectual horsepower and hunger to solve problems that don't have textbook answers yet, we want to talk to you., You will join a team chartered to make Apple's hyper-scale cloud network dramatically more fault-tolerant by building systems that detect and remediate failures automatically, with minimal or no human intervention. This means designing and building the sensing, decision-making, and remediation layers that allow the network to identify anomalies, localize root cause, and take corrective action in real time. You'll work on problems like: how do you detect a degrading network path before it causes an outage? How do you safely automate remediation actions on live production infrastructure without introducing new risk? How do you build a system that learns from past incidents to prevent recurrence? These are open, hard problems, and you will be expected to bring rigorous thinking, creativity, and strong engineering execution to solve them. This role is ideal for someone who loves taking a deep, formal understanding of distributed systems, control theory, graph algorithms, or fault-tolerant design patterns and turning it into resilient, real-world software running in production at enormous scale. You should be comfortable reading research papers and translating relevant ideas into working systems, while also being pragmatic about the constraints of operating in a live, hyper-scale environment. If you have exceptional technical depth from your academic background, competitive programming, research experience, or personal projects, and you're hungry to work on problems most companies aren't yet equipped to solve, this team is built for you. ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Concurrency with Go](https://www.wearedevelopers.com/videos/191-concurrency-with-go) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Characteristics of a Software Engineer](https://www.wearedevelopers.com/magazine/166-top-characteristics-of-a-software-engineer) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)