> Markdown version of [/jobs/ext/2874159-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2874159-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** Lumen Inc - **Location:** Hartford, CT, United States (Remote available) - **Experience:** Expert - **Salary:** $105,786.0 - $141,047.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Microsoft Azure, Cloud Computing, Continuous Integration, Distributed Systems, Ethernet, Fault Tolerance, Network Topologies, Virtual Private Networks (VPN), Python (Programming Language), Networking Basics, Systems Development Life Cycle, Reliability Engineering, Ansible, Prometheus, Software Engineering, Datadog, Google Cloud, Cloud Platform System, Computer Network Technologies, System Availability, Grafana, Multi-Cloud, Containerization, Kubernetes, Information Technology, Deployment Automation, Asynchronous Programming, Cloudwatch, Terraform, Open Network Automation Platform, Software Version Control - **Published:** September 13, 2026 - **Apply:** https://dejobs.org/x/x/D475F4AD3FA342C0B41136FB9F9A2824/job/ ## About the Role * Bachelor's degree or equivalent in engineering, computer science, or related field. * 8+ years in software development, systems engineering, and/or networking * 5+ years of related experience required. * Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP), including compute, networking, and identity services * Strong automation and infrastructure-as-code skills: Terraform, Ansible, and Python * Experience running containerized workloads on Kubernetes * Working knowledge of modern observability and monitoring tooling (e.g., Datadog, CloudWatch, Grafana, Prometheus), including building dashboards and defining alerts * Demonstrated experience with incident management and blameless postmortems * Comfort using AI-assisted development and agentic tools as part of daily engineering practice * Understanding of network technologies including Internet, Ethernet, IPVPN, Edge Compute, and Optical transport * Strong listening and communication skills; able to operate with autonomy while knowing when to escalate Preferred Qualifications: * Multi-cloud experience across AWS, Azure, and GCP * Asynchronous programming concepts and distributed systems design * Zero-downtime deployment strategies * High availability and multi-region architectures * Source control and CI/CD practices at scale * Experience applying agentic workflows to operational support and investigation ## Description Lumen's Network as a Service (NaaS) platform delivers on-demand networking at scale. As Lead SRE, you'll own the reliability of that platform - partnering with operations teams and development counterparts to drive technical direction and resolve systemic issues across a broad range of network topologies and applications. You'll be accountable for platform observability, incident management, and automation, and you'll coordinate across architecture, engineering, and systems development organizations to measurably improve reliability. You'll also use AI and agentic tooling to build utilities that accelerate deployment automation, platform administration, and incident investigation. Success in this role draws on networking fundamentals, cloud platforms, software development and troubleshooting methodology, and a bias toward automating what you'd otherwise do twice. We're looking for a change maker - someone who sees where the platform should go next and drives meaningful impact for the customers who rely on it, * Reliability & Observability * Serve as subject matter expert for network automation platform applications, services, and hosting environments * Build and maintain the observability stack: instrument services, collect and curate metrics, and create dashboards and visualizations that make system health obvious at a glance * Define and tune proactive alerting so issues surface before customers feel them * Champion core SRE principles - SLIs, SLOs, and error budgets - and advocate for resilient, fault tolerant architecture * Incident Management * Participate in an on-call rotation and lead incident response for service outages and unplanned downtime * Drive blameless postmortems and root cause analysis; own follow-up actions through to completion * Prevent recurrence through process improvements, tooling, and knowledge sharing across teams * Automation & Infrastructure * Automate deployment pipelines (CI/CD) and cloud infrastructure provisioning, scaling, and configuration using infrastructure as code * Develop tools and utilities that reduce toil and empower operations and development teams to manage services independently * Apply AI-assisted and agentic workflows to development, support, and investigation work * Collaboration & Leadership * Collaborate with cross-functional development teams to support, enhance, and scale NaaS applications * Provide guidance and mentorship to junior engineers * Maintain clear documentation for processes and architecture ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)