> Markdown version of [/jobs/ext/2272111-manager-infrastructure-sre-ai-platforms-services-special-projects](https://www.wearedevelopers.com/jobs/ext/2272111-manager-infrastructure-sre-ai-platforms-services-special-projects). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects - **Company:** Apple Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Databases, Data Centers, Data Infrastructure, Dataspaces, Disaster Recovery, Distributed Systems, PostgreSQL, Machine Learning, Open Source Technology, Redis, Reliability Engineering, AI Infrastructure, Google Cloud, Cloud Platform System, HybridCloud, AI Platforms, Kubernetes, Information Technology, Cassandra, Apache Kafka, Cloud Optimization - **Published:** August 27, 2026 - **Apply:** https://www.dice.com/job-detail/87d83b39-e2e1-40f9-80be-57b43ee19761 ## About the Role MS Degree in Computer Science or related degree and 12+ years of experience of progressive engineering leadership experience building, scaling, and operating mission-critical infrastructure platforms and global services. Management & Leadership Scope: 6+ years managing multi-layered engineering organizations (manager-of-managers) with a proven track record of hiring, developing, and retaining top-tier technical talent across global sites. Cloud & Distributed Compute Expertise: Demonstrated hands-on and architectural mastery of cloud-native infrastructure, Kubernetes platform engineering, and hybrid cloud operations (AWS, Google Cloud Platform, private data centers). Accelerated Computing & AI Infrastructure: Direct operational and architectural experience running large-scale systems for AI/ML training and inference workloads, including utilization optimization, scheduling, and high-performance storage/networking. SRE & Production Operations: Deep background in Site Reliability Engineering (SRE) principles, telemetry, observability frameworks, disaster recovery, and managing 24/7 high-availability infrastructure at scale. Technical Communication: Exceptional ability to seamlessly bridge executive strategy and low-level technical trade-offs-communicating vision to executive stakeholders while driving detailed technical discussions with principal engineers. Preferred Qualifications Large-Scale Enterprise Provenance: Experience leading core infrastructure or foundational platform SRE for a global, tier-1 technology organization operating at massive scale. Multi-Engine Database & Data Infrastructure: Familiarity overseeing diverse open-source and proprietary storage/data ecosystems (e.g., Cassandra, FoundationDB, Kafka, Redis, PostgreSQL). Financial & Capacity Governance: Proven competency managing large-scale infrastructure investments, capital expenditures, operational budgets, capacity forecasting, and cloud optimization strategies. ## Description We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long-term technical strategy, organizational structure, and operational roadmap for global, mission-critical infrastructure platforms on the Services Special Projects team. This position requires a rare blend of deep technical domain expertise-spanning distributed systems, Kubernetes, and AI workload orchestration-and proven organizational leadership managing large, globally distributed engineering teams., In this role, you will be responsible for defining and building infrastructure strategy that balances continuous innovation with high reliability, performance, and cost efficiency. You will lead a growing, multi-tiered team of engineers who are responsible for foundational platforms that power large-scale consumer and enterprise workloads. Beyond operational delivery, you will establish standards for operational excellence, Site Reliability Engineering (SRE), and capacity planning. You will be a key strategic partner, translating complex business imperatives into scalable platform designs while cultivating a strong engineering culture focused on automation, technical ownership, accountability, and continuous improvement. ## Related Videos - [Maximising Cassandra's Potential: Tips on Schema, Queries, Parallel Access, and Reactive Programming](https://www.wearedevelopers.com/videos/1167-maximising-cassandra-s-potential-tips-on-schema-queries-parallel-access-and-reactive-programming) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Reference Architecture of AI in the Cloud](https://www.wearedevelopers.com/videos/1613-reference-architecture-of-ai-in-the-cloud) - [Building Real-Time AI/ML Agents with Distributed Data using Apache Cassandra and Astra DB](https://www.wearedevelopers.com/videos/782-building-real-time-ai-ml-agents-with-distributed-data-using-apache-cassandra-and-astra-db) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)