> Markdown version of [/jobs/ext/97893-site-reliability-engineer-human-engineering](https://www.wearedevelopers.com/jobs/ext/97893-site-reliability-engineer-human-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Human Engineering - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Salary:** $181,100.0 - $318,400.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Systems Engineering, Continuous Integration, Distributed Systems, Django Web Framework, Domain Name System (DNS), Elasticsearch, Identity and Access Management, Python (Programming Language), PostgreSQL, Enterprise Messaging Systems, Networking Basics, Open Source Technology, OpenID, RabbitMQ, Redis, Reliability Engineering, Software Tools, Service Discovery, Web Applications, Workflow Management Systems, YAML, Datadog, SSL Certificate Management, Data Logging, Transport Layer Security, Load Balancing, Istio, Amazon Virtual Private Cloud (VPC), Backend, Event Driven Architecture, Amazon Relational Database Service, Kubernetes, Information Technology, Apache Kafka, Celery, Api Gateway, Firewall Services Module, Terraform, Data Pipelines, Dynatrace - **Published:** May 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e635c3bc13c6c46f ## About the Role Do you have experience in Web applications?, We're looking for a Site Reliability Engineer who thinks like a systems engineer first and an operator second. You won't just keep things running - you'll shape how our platform evolves. Our team operates 50+ services across Kubernetes and AWS, handles sensitive health and research data, and is ramping up many architectural shifts: new service-to-service auth patterns, event-driven pipelines, and a move from on-prem to cloud-native infrastructure. We need someone who gets excited about that kind of work, can reason about distributed systems at the design level, and is a strong enough communicator to bring the rest of the team along., BS in Computer Science, Engineering, or equivalent practical experience, with 7+ years of experience in distributed systems Experience with event-driven architectures (Kafka, RabbitMQ, or similar messaging systems) Experience with service mesh or API gateway patterns (Istio, Envoy, Kong, or similar) Familiarity with Django/Python web applications and their operational characteristics (Celery, Gunicorn, PostgreSQL) Experience with observability tooling beyond basic monitoring: distributed tracing, SLO frameworks, structured logging Background working with sensitive data (health data, PII) and associated compliance requirements Experience leading incident response and building on-call culture Contributions to internal or open-source infrastructure tooling, BS in Computer Science, Engineering, or equivalent practical experience, with 5+ years of experience in distributed systems Deep experience with Kubernetes in production - cluster operations, networking, storage, troubleshooting Strong proficiency designing and operating services in AWS (EC2, EKS, RDS, S3, IAM, VPC) Hands-on infrastructure-as-code experience (Terraform, Helm, or equivalent) Proficiency in at least one backend language (Python, Go, or similar) - you can write production services, not just scripts Experience with CI/CD pipeline design and GitOps workflows Strong understanding of networking fundamentals: DNS, load balancing, TLS, firewall rules, service discovery Excellent communication skills. You can explain a complex system to a room of engineers who didn't build it Experience building internal automation or self-service tooling (Slack bots, CLI tools, workflow orchestration) that reduced manual operational work ## Description The Human Engineering Software team builds tools used across Apple for user studies, research participant management, health data collection, and privacy-preserving analytics. Our infrastructure spans Django backends, Kubernetes clusters (self-hosted and AWS), PostgreSQL, Redis, Kafka, Elasticsearch and a growing set of internal service integrations. This role is engineering-forward SRE. You'll spend as much time designing systems as operating them. You'll work closely with our full-stack engineers to improve how services communicate, how we observe production behavior, and how we ship changes safely. You'll have a seat at the architecture table - we want you proposing solutions, not just implementing them. ","responsibilities":"Platform & Reliability Engineering - Own the reliability of our Kubernetes-hosted services across AWS and self-hosted clusters: deployments, scaling, capacity planning, certificate management, and secrets rotation. Design and implement SLO-driven observability: define meaningful SLIs, build dashboards that answer "is the system healthy?" not just "is the pod running?" Drive incident response and blameless postmortems Distributed Systems & Architecture - Partner with the architecture team on system design: service-to-service authentication (OIDC, gateway auth), event-driven messaging (Kafka), API gateway patterns. Design the infrastructure layer to make architecture proposals real in production. Evaluate and recommend new tools, patterns, and platforms and write code when it's the right tool, whether that's a deployment operator, a health check service, or a data pipeline component. This isn't a YAML-only role Engineering Enablement - Make the team efficient; own CI/CD pipelines and GitOps practices, owning tests to verify or production tools are functioning correctly, build self-service automation, evolve our observability and security posture, and communicate infrastructure decisions clearly across technical and non-technical stakeholders ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries) - [Dev Digest 119 - ❤️ === ❤️](https://www.wearedevelopers.com/magazine/454-dev-digest-119)