Site Reliability Engineer

Apple Inc.
Austin, TX, United States
3 days ago
Apply on www.themuse.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Cloud Computing Computer Networks Databases Linux DevOps Distributed Data Store Distributed Systems Domain Name System (DNS) Monitoring of Systems Python (Programming Language) Oracle (Applications)
+17 more
Reliability Engineering Software Engineering Apache Solr Transport Layer Security Enterprise Software Applications Load Balancing Containerization Information Technology Cassandra Apache Kafka Build Tools Hardware Infrastructure Splunk Docker Golang Programming Languages Microservices

Job description

As an SRE, you will play a key role in ensuring the reliability, scalability, and performance of Apple’s FMD (Full Material Disclosure) Platform while driving automation and DevOps best practices.

Responsibilities:

System Observability: Implement and maintain robust observability solutions that provide real-time insights into system performance and health, proactively identifying and addressing potential issues before they impact users.

Troubleshooting and Root Cause Analysis: Investigate and resolve incidents swiftly during critical situations, performing thorough root cause analysis to prevent recurrence.

Automation: Leverage your coding skills to build tools and automate runbooks, driving efficiency and reducing manual toil.

Documentation: Maintain and manage runbooks and best practices to foster knowledge sharing and improve team efficiency.

Requirements

Communication: Demonstrate strong interpersonal skills with the ability to collaborate effectively across multiple business and technical teams.

Preferred Qualifications

Good understanding of database principles and working knowledge in distributed storage and infrastructure solutions such as Oracle, Cassandra, SOLR, and Kafka (1-2 yrs of experience).

Good command on Linux, networking concepts (TLS/SSL, DNS, Load Balancers, etc.,) and troubleshooting skills in large scale environments (1-2 yrs of experience).

Experience with container management and microservices architectures such as Docker, across cloud and on-premises infrastructure.

Minimum Qualifications

3+ years of experience in Site Reliability Engineering, DevOps, Software Engineering, or a related field in the US.

Proficiency in at least one scripting or programming language (Python, Go, Java, or similar).

Monitoring distributed system architectures with hands-on expertise in log monitoring and analysis tools such as Splunk.

Education: Bachelor’s or Master’s degree in Computer Science or a related field (equivalent practical experience).

About the company

Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, or Apple Store experience we deliver is the result of us making each other’s ideas stronger. That happens because every one of us shares a belief that we can make something wonderful and share it with the world, changing lives for the better. It’s the diversity of our people and their thinking that inspires the innovation that runs through everything we do. When we bring everybody in, we can do the best work of our lives. Here, you’ll do more than join something - you’ll add something.”

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.themuse.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

Videos

See all

Related articles

See all