> Markdown version of [/jobs/ext/248730-site-reliability-engineer-physical-infrastructure](https://www.wearedevelopers.com/jobs/ext/248730-site-reliability-engineer-physical-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, Physical Infrastructure - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Experienced - **Salary:** $172,100.0 - $258,600.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Unix, Command-Line Interface, Data Centers, DevOps, Domain Name System (DNS), Hypertext Transfer Protocols (HTTP), Python (Programming Language), Reliability Engineering, Shell Script, Software Engineering, TCP/IP, Data Logging, Diagnostic Tools, System Availability, Large Language Models, Grafana, Reliability of Systems, Technical Debt, Generative AI, Kubernetes, Information Technology, Performance Monitor, Docker - **Published:** May 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=597969fc36bdf870 ## About the Role Do you have experience in System performance monitoring?, Do you have a Bachelor's degree?, We are looking for a creative and highly motivated Site Reliability Engineer to join our team. Having depth and breadth of knowledge working in physical infrastructure in a large-scale distributed environment is a strength you'll need. You should have experience in unix systems administration, DevOps, and data center infrastructure. If you are passionate about solving complex problems at scale, we want to hear from you!, Build automation tools that eliminate routine tasks. Every manual process is an opportunity to code a solution Experience with Unix/Linux systems administration and command-line diagnostic tools Proven experience leading initiatives to reduce technical debt, refactor systems, or improve performance and latency Expertise in performance analysis and capacity planning for physical infrastructure. Demonstrated ability to lead incident response for high-impact outages Familiarity with using Generative AI (GenAI) or Large Language Models (LLMs) to accelerate operational tasks, such as automating runbooks, generating scripts, or analyzing incident data Minimum Qualifications 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Systems Admin focused on physical infrastructure in a large-scale distributed environment Strong software development skills in a language like Swift, Go, or Python, and a high degree of comfort with shell scripting (Bash) Hands-on experience building and managing systems with container orchestration tools (Kubernetes, Docker) Deep understanding of networking (TCP/IP, DNS, HTTP) and experience using observability tools (monitoring, logging, tracing) to diagnose complex issues Excellent problem-solving and communication skills, with a strong sense of ownership and drive BS/MS in Computer Science, Engineering or related field ## Description The Systems and Infrastructure team builds and manages world class services and physical infrastructure for Apple software engineers world wide to build, test, and release Apple's software. About Our Team: We are a team dedicated to engineering excellence, reusable design, and simplicity. We foster a supportive, growth-focused culture where we mentor each other and work together to build resilient, high-quality systems. ","responsibilities":"Ensure System Reliability: Design, build, and maintain robust, scalable, and observable systems for our core infrastructure services Automate: Reduce operational toil by developing automation and tooling to prevent and rapidly resolve production issues Improve Incident Response: Own and refine our incident management processes to ensure high availability Collaborate with Engineers: Partner with development teams to create functional, high-quality solutions that support the entire workflow Improve and Modernize Systems: Use a proactive approach to identify and eliminate technical debt to enhance long-term reliability and maintainability ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 119 - ❤️ === ❤️](https://www.wearedevelopers.com/magazine/454-dev-digest-119) - [ Dev Digest 213: Petrol Prices, Agentic Workflows, AI Skills and CODE100!](https://www.wearedevelopers.com/magazine/718-dev-digest-213-petrol-prices-agentic-workflows-ai-skills-and-code100)