> Markdown version of [/jobs/ext/1055705-principal-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1055705-principal-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - **Company:** Zayo Group LLC - **Location:** Denver, CO, United States - **Experience:** Expert - **Salary:** $116,700.0 - $179,500.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Application Performance Management, Border Gateway Protocol, Computer Networks, Linux, Domain Name System (DNS), Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Nagios, Netconf, Reliability Engineering, Ansible, Prometheus, TCP/IP, Scripting, Transport Layer Security, Grafana, Juniper, Kubernetes, Information Technology, Cacti, Puppet, Cisco, Docker - **Published:** June 30, 2026 - **Apply:** https://diversityjobs.com/career/17431039/Principal-Site-Reliability-Engineer-Colorado-Denver ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience). * Minimum of ten (10) years of experience in a Site Reliability Engineering or related role. * Strong understanding of system administration, Linux, and scripting languages (Python and various shells). * Expert at developing automation tools for monitoring, alerting, and deployment to ensure efficient and reliable operations. * Expert at designing and implementing monitoring systems at scale. * Expert at container orchestration (Kubernetes and Docker). * Experience with monitoring platforms such as SevOne, Assure1, and Nagios and various vendor NMS systems. * Previous work in large scale distributed production environments. * Experience with a variety of cloud platforms and tools (AWS, Google, etc). * Experience with a variety of monitoring and alerting tools (Prometheus, Grafana, Cacti, etc.) * Strong working knowledge of networking concepts and application protocols, especially TCP/IP, BGP, DNS, TLS, and HTTP/S. * Experience with infrastructure management tools such as Ansible, Terrafor, Puppet, to deploy and manage infrastructure at scale. * Proven leadership skills, with the ability to mentor and inspire others. * Excellent problem-solving, analytical, and critical thinking skills. * A passion for automation and building efficient systems. Bonus Points if you have: * Experience working with various vendor APIs (or netconf) including Nokia, Juniper, Fujitsu, Infinera, Cisco, and Ciena. * Experience with various network orchestration platforms such as Ciena Blue Planet MDSO, Cisco NSO, Nokia NSP, or others. ## Description * Automation: Develop and implement automation solutions to streamline operations and reduce manual effort. * Monitoring and Alerting: Design and implement effective monitoring and alerting systems to proactively identify and address issues. * Incident Management: Own the incident lifecycle, from leading root cause analysis and resolution to implementing preventative measures to avoid future occurrences. Be on-call to diagnose and resolve critical service outages. * Reliability Engineering: Proactively identify and mitigate potential system risks, focusing on automation, monitoring, and tooling to ensure high service availability. * Scalability and Performance: Design and implement solutions to ensure our infrastructure can handle ever-growing demands while maintaining optimal application performance. * Collaboration: Work closely with developers, product managers, and other engineers to translate business needs into robust and reliable technical solutions. Become the beacon for best practices and efficient processes throughout the organization. ## Related Videos - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Unlocking the potential of Digital & IT at Vodafone](https://www.wearedevelopers.com/videos/602-unlocking-the-potential-of-digital-it-at-vodafone) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)