Apple Service Engineering - Data Streaming SRE

Apple Inc.
Washington, United States
14 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Systems Engineering Data as a Services Data Centers Data Infrastructure Database Storage Structures DevOps Distributed Systems Fault Tolerance Internet Services Python (Programming Language)
+11 more
Open Source Technology Reliability Engineering Software Engineering Data Streaming Google Cloud Cloud Platform System Kubernetes Infrastructure Automation Frameworks Apache Kafka Terraform Golang

Job description

The Data Service SRE team develops applications and tooling that are safe, reliable, scalable, and fast. This work requires an innovative spirit and an extraordinary degree of care and difficulty in engineering. Team members contribute to all major components of Kafka deployment infrastructure, including maintenance automation, control plane enhancements, monitoring and alerting tooling/dashboards, advanced deployment architecture, focused on safety, stability, performance, and scaling.

Requirements

The Apple Service Engineering - Data Streaming SRE team is looking for Site Reliability Engineers with experience developing processes, tools, and automation for managing distributed systems in production environments. Our SRE team combines software engineering, systems engineering, and Devops practices to build and run large-scale, massively distributed, fault-tolerant systems. Our software ensures that Apple’s services are reliable, scalable, and secure, and we leverage both open-source and homegrown technologies to provide managed data infrastructure services. You will help build next-generation Kafka infrastructure and platform services, collaborating cross-functionally with various ASE teams-from store and commerce to search and recommendations. You’ll create platforms that can rapidly scale to serve data with very low latencies. You should be someone who isn’t afraid to question assumptions, thrives as a collaborative partner under tight deadlines, and tackles complex problems with elegant technical solutions., 5 or more years of experience in support of internet-facing production services and distributed systems via deployments, On Call and Incident Management.

5 or more years of experience running large scale infrastructure with a heavy reliance on automation tooling

5 or more years of experience troubleshooting and performance deep dive analysis

Real operational experience managing services at scale on Kubernetes

Proficient in one or more of the following programming languages: Java, Go (golang), Python

Operational experience deploying in and running on Datacenter and Cloud architectures (networking topologies, host placement strategies, and failure modes); design of multi-datacenter systems; failure domains; and wide-area networking.

Self motivated, inquisitive with an aptitude to learn new technologies quickly and effectively.

Demonstrated expertise developing and troubleshooting distributed systems and database storage engines.

Experience developing critical internet services and/or platform infrastructure.

Experience with AWS, Google Cloud Platform and IaC such as Terraform

Preferred Qualifications

Experience managing messaging services such as Kafka or other Data services

Proficient in Java, Go (golang) & Python

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

2:38 min

Extreme engineering culture for massive web data operations

Ariel Shulman Ariel Shulman +1 · WWC Europe 2026

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

Videos

See all

Related articles

See all