> Markdown version of [/jobs/ext/2077431-data-platform-sre-ai-data-platforms-aidp](https://www.wearedevelopers.com/jobs/ext/2077431-data-platform-sre-ai-data-platforms-aidp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Platform SRE, AI & Data Platforms (AiDP) - **Company:** Apple Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Artificial Intelligence, Airflow, Amazon Web Services, Apache HTTP Server, Big Data, Command-Line Interface, Cloud Computing, Computer Programming, Data Infrastructure, Data Warehousing, Software Debugging, Distributed Computing Environment, Distributed Systems, Fault Tolerance, Python (Programming Language), Open Source Technology, Performance Tuning, Systems Development Life Cycle, Scala (Programming Language), Secure Coding, Software Construction, Software Engineering, Graphics Processing Unit (GPU), System Availability, Large Language Models, Apache Spark, Data Lakes, Kubernetes, Low Latency, Free and Open-Source Software, Data Management, Machine Learning Operations, Data Pipelines - **Published:** August 16, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3355967026&tx=DT110UYI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * 3+ years of professional software engineering experience with large-scale big data platforms, including strong programming skills in Java, Scala, Python, or Go. * Proven expertise in operating large-scale distributed data processing systems with a strong focus on Apache Spark. * Hands-on experience with table formats and data lake technologies such as Apache Iceberg, ensuring scalability, reliability, and optimized query performance. * Strong background in incident management, including troubleshooting, root cause analysis, and performance optimization in complex production environments. * Proficient with cloud technologies such as AWS and GCP * Experience with Unix/Linux systems and command-line tools for debugging and operational support., * Expertise in designing, building, and operating critical, large-scale distributed systems with a focus on low latency, fault-tolerance, and high availability. * Experience with contribution to Open Source projects is a plus. * Experience with multiple public cloud infrastructure, managing multi-tenant Kubernetes clusters at scale and debugging Kubernetes/Spark issues. * Experience with workflow and data pipeline orchestration tools (e.g., Airflow, DBT). * Understanding of data modeling and data warehousing concepts. * Familiarity with the AI/ML stack, including GPUs, MLFlow, or Large Language Models (LLMs). * A learning attitude to continuously improve the self, team, and the organization. * Solid understanding of software engineering best practices, including the full development lifecycle, secure coding, and experience building reusable frameworks or libraries. ## Description As a Data Platform SRE, you will be responsible for developing and operating our big data platform using open source or other solutions to aid critical applications, such as analytics, reporting, and AI/ML apps. This includes working to optimize performance and cost, automate operations, and identifying and resolving production errors and issues to ensure the best data platform experience. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)