> Markdown version of [/jobs/ext/1069969-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1069969-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Thunderbird LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $123,000.0 - $144,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Continuous Integration, Software Debugging, Github, Identity and Access Management, Internet Message Access Protocols, Key Management, Simple Mail Transfer Protocols, Network Segmentation, Open Source Technology, OpenID, Reliability Engineering, Security Assertion Markup Language (SAML), Service Design, Software Engineering, Pulumi, Okta, Grafana, Session Description Protocol Security Descriptions (SDES), Kubernetes, Free and Open-Source Software, Terraform - **Published:** June 9, 2026 - **Apply:** https://www.builtincolorado.com/job/senior-site-reliability-engineer/9649698 ## About the Role * 7+ years of experience in infrastructure, platform engineering, or site reliability roles, including hands-on production Kubernetes experience in workload operations, troubleshooting, and cluster management. * Hands-on experience with infrastructure-as-code on AWS using Terraform, OpenTofu, or Pulumi. * Security awareness in day-to-day infrastructure work: identity, least privilege, secrets hygiene, and network controls. * Demonstrated ownership mindset with the ability to proactively identify issues, drive work to completion, and communicate risks early. * Excellent async written communication skills; comfortable working with a geographically distributed team. * Ability to collaborate effectively with software engineers and non-engineering stakeholders to improve platform reliability and operational efficiency. * Ability to learn, evaluate, and responsibly use emerging technologies, including AI-enabled tools, to improve work processes. Bonus points for * Experience with GitOps workflows (ArgoCD or Flux). * Familiarity with Keycloak or similar identity platforms (OIDC, SAML, federation). * Knowledge of email protocols and/or experience operating email infrastructure (SMTP, IMAP). * Prior work in or alongside an open-source community. * French, German, Japanese, or other language proficiency in addition to English. ## Description The Senior Site Reliability Engineer establishes and maintains the infrastructure and operational systems that Thunderbird users and teams depend on every day. You'll design and develop CI/CD systems for MZLA websites, services, and release workflows, diagnose and debug production incidents, and implement improvements to enhance system reliability. We believe that good infrastructure work is invisible when it's going well and invaluable when it isn't. This role is for someone who treats production as something to be understood, not just kept running. You write things down, flag problems before they become fires, and leave documentation better than you found it. You bring production instincts, infrastructure-as-code fluency, and security awareness that's baked in, not bolted on. You'll work closely with Software Development Engineers, team members, and community contributors, reporting to the Sr Manager, Platform Infrastructure. This is a great opportunity for someone who thrives with ambiguity, makes good decisions without a complete picture, and cares about Thunderbird's mission: open-source software used by millions who choose privacy and ownership over convenience. This role requires consistent overlap with Pacific Time zone working hours to enable effective collaboration. You should have availability for regular overlap hours for context sharing with Pacific Time colleagues. What you'll do * Operate and evolve our EKS-based Kubernetes platform, supporting service migrations, platform improvements, and reliability initiatives. * Design and develop CI/CD systems supporting websites, services, and Thunderbird desktop releases, contributing to pipeline reliability and OIDC-based authentication across GitHub Actions workflows. * Write and maintain infrastructure in Pulumi and/or Terraform/OpenTofu across multiple AWS accounts. * Operate and evolve our observability stack (VictoriaMetrics, VictoriaLogs, Grafana, Vector) and partner with engineering teams to incorporate instrumentation and monitoring into service design. * Apply security-conscious infrastructure practices, including least-privilege IAM, secrets management via AWS Secrets Manager and External Secrets Operator, and network segmentation. * Diagnose and debug production incidents; drive root-cause analysis and post-incident improvements to prevent recurring problems. * Participate in on-call rotation and collaborate with SDEs and fellow SREs to ship, maintain, and monitor new builds and support service onboarding. * Contribute to runbooks, architecture documentation, and team processes. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)