Site Reliability Engineer

Apple Inc.
Austin, TX, United States
1 day ago
Apply on www.themuse.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Amazon S3 Build Automation Cloud Computing Databases Continuous Integration Couchbase Servers Domain Name System (DNS) Hypertext Transfer Protocols (HTTP) Python (Programming Language) MongoDB Network Protocols Reliability Engineering
+16 more
Ansible Prometheus Shell Script Software Engineering Transmission Control Protocol (TCP) Delivery Pipeline Grafana Mttr Generative AI Kubernetes Storage Technologies Information Technology Low Latency Cassandra Hardware Infrastructure Splunk

Job description

The Customer Systems Team is looking for an experienced Site Reliability Engineer. In this role you will design, build and deliver highly scalable, reliable, secure cloud infrastructure which powers the applications and services used by Apple’s customers every day. You will work closely with cross functional teams, business leaders and other partners across Apple to implement new solutions. If infrastructure as code, automation and intelligent monitoring excites you then this is the job for you., Innovate, architect, build, and document highly available, scalable, reliable, secure Infrastructure

Troubleshoot application specific, network, system & performance issues

Build and maintain CI/CD infrastructure to enable fast delivery cycles for software engineering teams

Envision and build automation tools to deliver infrastructure services reliably and in a repeatable fashion

Collaborate with other site reliability engineers, software engineers, quality engineers, to gather, define, and analyze non-functional/technical requirements

Requirements

Excellent problem solving, critical thinking, and interpersonal skills

Good communication skills to collaborate with distributed teams

Experience with Cassandra, MongoDB, Couchbase databases, AWS S3 or similar storage technologies

Experience in deploying, monitoring and supporting java applications

Experience with ArgoCD and GitOps model

Experience in defining, monitoring and achieving key operational metrics like MTTR and SLO

Experience with GenAI tools in workflow automation for infrastructure management

Ability to learn new technologies in a short time

Minimum Qualifications

5+ years of experience in designing and building resilient, large-scale, low latency, cloud and on-prem Infrastructure including Compute, Storage, and Network

3+ years of experience with deploying/managing Kubernetes using Helm

Experience with Shell Scripting, Python, or Ansible

Experience in monitoring using Splunk, Grafana, Prometheus, Alertmanager

Deep understanding of networking protocols: DNS, TCP, HTTP/HTTPS

Experience in setting up and managing CI/CD pipelines

Bachelor’s or Master’s in Computer Science or equivalent experience

About the company

Imagine what you could do here. Apple is a place where extraordinary people gather to do their best work. Together we craft products and experiences people once couldn’t have imagined - and now can’t imagine living without. If you’re motivated by the idea of making a real impact, and joining a team where we pride ourselves in being one of the most diverse and inclusive companies in the world, we’d love to hear from you!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.themuse.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:36 min

Critical infrastructure and performance skills for modern developers

Andrew Holway · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all