Sr. Linux Engineer - DNS and Global Server Load Balancing, Infrastructure Services

Apple Inc.
Sunnyvale, CA, United States
23 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$216,200.0
Working hours
Regular working hours

Tech stack

Build Automation Border Gateway Protocol Configuration Management Code Review Continuous Integration Data Centers Linux Disaster Recovery Distributed Systems Domain Name System Security Extensions Domain Name System (DNS) Python (Programming Language)
+21 more
Key Management Network Layer Linux Kernel Networking Basics Routing Network Service Package Management Systems Ansible Prometheus Software Deployment Datadog Load Balancing Computer Network Operations Cloud Platform System Performance Testing Grafana Information Technology Bare Metal Puppet Terraform Programming Languages

Job description

Infrastructure Services is part of IS&T and the foundation of Apple’s global network operations - managing data center equipment and systems to deliver compute, storage, and networking services for teams across Apple, including its internal developer community. From individual facilities to a worldwide network, Infrastructure Services ensures the technology underneath everything works without question

Apple’s services reach hundreds of millions of people every day, and Edge Services builds and runs the authoritative DNS and global server load balancing infrastructure that sits at the front door of all of them. Every request to an Apple service begins here, in the first few milliseconds where our systems decide how and where to answer. The reliability, correctness, and performance of that front door directly shape the experience people have with Apple, and keeping it fast and dependable at global scale is the work of our team.

We’re looking for a thoughtful, experienced engineer to bring technical leadership to this infrastructure. You’ll operate and extend bare-metal and cloud systems distributed across dozens of data centers and network points of presence worldwide, build deep traffic and resolution visibility, drive performance and capacity planning, and create the automation that keeps our 24x7 operations calm and reliable. You’ll also mentor engineers across the team and partner closely with our software, network, and product engineering groups. We believe great systems are built by great teams and care as much about how we work together as about what we ship. If you’re energized by owning the reliability of critical infrastructure at global scale, we’d like to hear from you., This is a senior individual-contributor role and a technical anchor for the Edge Services DNS and GSLB systems. The essential functions of the job are to operate and extend distributed Linux infrastructure as code across geographically dispersed sites; to define and maintain the observability, alerting, and service-level objectives that keep the systems healthy; to build automation that removes operational toil; to plan fleet capacity and hardware lifecycle across globally distributed sites; and to mentor engineers while partnering across software, network, and product teams. The role includes incident and process management and participation in a shared 24x7 production on-call rotation.

Responsibilities

Operates and extends bare-metal and cloud Linux infrastructure across dozens of edge sites and points of presence, treating infrastructure as code so that standing up or rebuilding a site is predictable and repeatable.

Owns the reliability, correctness, and performance of the authoritative DNS and global server load balancing systems, including incident response and participation in a 24x7 on-call rotation.

Builds deep visibility into the systems by defining meaningful service-level objectives, alerting, and dashboards that span every point of presence and surface problems early.

Develops automation and tooling in Python, Go, Rust, and/or Swift - supported by generative AI tooling - to remove operational toil and turn manual tasks into reliable, self-sustaining systems.

Guides fleet capacity, growth, and hardware lifecycle across geographically distributed sites through smart tooling, thoughtful planning, and collaboration with leadership and program management.

Mentors and grows engineers at every career stage and partners with other senior engineers to raise the bar on architecture, tooling, and design.

Partners with software, network, and product engineering teams to deliver large-scale rollouts and keep 24x7 operations running smoothly.

Represents the team’s work, needs, and ideas to partners and leadership through clear, collaborative communication.

Requirements

Demonstrated experience operating Linux systems in production, including software deployment and CI/CD workflows.

Working knowledge of networking fundamentals, with hands-on experience troubleshooting TCP/UDP and common layer 2-3 issues.

Experience with configuration management or infrastructure-as-code tooling (for example, Salt, Ansible, Puppet, or Terraform) and with observability tooling (for example, Prometheus and Grafana, or equivalents).

Proficiency in at least one programming language used for automation and tooling (for example, Python, Go, Rust, or Swift).

Willingness to participate in a shared 24x7 on-call rotation.

Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.

Preferred Qualifications

Extensive experience operating large-scale infrastructure across multiple data centers or network points of presence.

Expertise with anycast routing and BGP, and a strong understanding of how DNS, load balancing, and the network layer interact.

Depth in Linux internals, including kernel networking, and in package management and software deployment at fleet scale.

Experience defining service-level objectives and indicators (SLOs/SLIs) and driving reliability programs across distributed systems.

Familiarity with secure-by-default operations, such as DNSSEC key management, change safety, and progressive configuration rollout.

Experience with scale and performance testing, disaster recovery, and capacity planning.

Experience managing hardware lifecycle across edge and regional network environments.

Track record of leading end-to-end projects, defining technical roadmaps, and driving cross-functional alignment on architecture and best practices.

Experience mentoring engineers and leading code reviews.

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $216,200 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:26 min

Understanding Puppeteer and its underlying architectural design

Miki Lombardi · JS Congress

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all