> Markdown version of [/jobs/ext/3040589-director-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/3040589-director-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director, Site Reliability Engineering - **Company:** DUCKDUCKGO SUBSCRIPTION INC. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $244,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Artificial Intelligence, Android Software Development, Macintosh Computers, Apple IOS, Configuration Management, DuckDuckGo (Internet Search Engines), Linux, Distributed Systems, Software Deployment, Software Engineering, Maintaining Code, Search Engines, Web Technologies, GPT, Serverless Computing, Docker - **Published:** September 23, 2026 - **Apply:** https://www.builtincolorado.com/auth/login?destination=/job/director-site-reliability-engineering/11311270 ## About the Role * 10+ years relevant professional experience in reliability, platform, infrastructure, or software engineering, including 4+ years leading SRE teams. * Experience participating in a 24x7 on-call rotation for a large-scale deployment. * Ability to lead and collaborate on high-impact and complex projects from proposal through postmortem. * Proficient in AI-driven development, including designing and implementing agentic workflows * Skills to wrangle vague problems, propose innovative solutions, and execute them with a strong focus on metrics. * Experience developing effective tools, services, alerts, and responses to identify and address reliability risks. * Investigative ability to root-cause sources of instability in high-traffic, distributed systems. * Deep experience administering and troubleshooting Linux and web technologies. * Ability to implement automation around infrastructure provisioning and configuration management to prioritize efficiency, scalability, and reliability. * Foresight to help identify the future technical direction of our deployment with the goal of improving reliability and performance. * Advanced programming skills enabling close partnership with software engineers to triage production issues and identify appropriate remediation, including code changes and performance considerations. * Ability to leverage cloud-native services and architectures to enhance reliability and scalability, with hands-on experience packaging and deploying applications using Docker and Docker Compose. ## Description Leads site reliability engineering for large-scale, privacy-focused infrastructure. Responsibilities include managing SRE teams, solving complex software and systems reliability challenges, supporting 24x7 operations, troubleshooting distributed Linux and web systems, developing automation, improving observability and incident response, deploying applications with Docker, and shaping technical direction. The role also requires advanced programming, AI-driven development, agentic workflows, and ownership of high-impact projects from proposal through postmortem., Working on the Site Reliability Team, you'll help build and maintain world-class infrastructure to meet the needs of millions of users protecting their privacy online. You'll utilize high-level languages like Perl, Go, TypeScript, or Python and work on related projects. Recent projects include: * Ensuring our Duck.ai product meets our reliability standards and minimizing user friction on failures * Scaling up our own index infrastructure to handle billions of documents * Create anti fraud verifications that respect users privacy As Director, Site Reliability Engineering, you'll dive deep into complex operational challenges, including software, systems, automation, and process analysis. We are looking for candidates who can read, write, troubleshoot, and deploy all types of software to help us tackle the reliability challenges of large-scale deployments., We want to ensure that our hiring process is accessible. If you need reasonable accommodation for any part of the application process because of a medical condition or disability, please send an email to [email protected] to let us know the nature of your request. Please note that: * You'll be required to attend meetings on camera via video conferencing * Expect to travel at least two times a year: once for our all-hands meetup and again for a team retreat (each around 4-5 days). While extenuating circumstances may impact attendance, everyone is strongly encouraged to attend. * While we offer a flexible work arrangement with no core hours, expect an average full-time commitment of 40 hours per week. * A successful candidate must pass a background check as a condition of joining the team. * By applying for this role, you confirm that all information submitted is accurate and complete. You further acknowledge that providing false or fraudulent information during the application process is cause for denial of an offer, revocation of any existing offer, or other adverse action, up to and including termination after the start of your commencement of work. Disclosure Statement: Use of AI in Hiring Process As part of our commitment to enhancing our recruitment process, we utilize artificial intelligence (AI) technology to assist in reviewing and summarizing job applications and test projects, including those tools integrated into our recruitment vendor platforms. We use AI to flag potentially fraudulent applications, analyze and summarize applicants' experience, interviews, and project performance, and help streamline our selection process. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)