Lead Site Reliability Engineer

Lumen Inc
Boston, MA, United States
8 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$105,786.0 - $141,047.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Systems Engineering Microsoft Azure Cloud Computing Continuous Integration Distributed Systems Ethernet Fault Tolerance Network Topologies Virtual Private Networks (VPN) Python (Programming Language)
+22 more
Networking Basics Systems Development Life Cycle Reliability Engineering Ansible Prometheus Software Engineering Datadog Google Cloud Cloud Platform System Computer Network Technologies System Availability Grafana Multi-Cloud Containerization Kubernetes Information Technology Deployment Automation Asynchronous Programming Cloudwatch Terraform Open Network Automation Platform Software Version Control

Job description

Lumen’s Network as a Service (NaaS) platform delivers on-demand networking at scale. As Lead SRE, you’ll own the reliability of that platform - partnering with operations teams and development counterparts to drive technical direction and resolve systemic issues across a broad range of network topologies and applications.

You’ll be accountable for platform observability, incident management, and automation, and you’ll coordinate across architecture, engineering, and systems development organizations to measurably improve reliability. You’ll also use AI and agentic tooling to build utilities that accelerate deployment automation, platform administration, and incident investigation.

Success in this role draws on networking fundamentals, cloud platforms, software development and troubleshooting methodology, and a bias toward automating what you’d otherwise do twice. We’re looking for a change maker - someone who sees where the platform should go next and drives meaningful impact for the customers who rely on it, * Reliability & Observability

  • Serve as subject matter expert for network automation platform applications, services, and hosting environments
  • Build and maintain the observability stack: instrument services, collect and curate metrics, and create dashboards and visualizations that make system health obvious at a glance
  • Define and tune proactive alerting so issues surface before customers feel them
  • Champion core SRE principles - SLIs, SLOs, and error budgets - and advocate for resilient, fault tolerant architecture
  • Incident Management
  • Participate in an on-call rotation and lead incident response for service outages and unplanned downtime
  • Drive blameless postmortems and root cause analysis; own follow-up actions through to completion
  • Prevent recurrence through process improvements, tooling, and knowledge sharing across teams
  • Automation & Infrastructure
  • Automate deployment pipelines (CI/CD) and cloud infrastructure provisioning, scaling, and configuration using infrastructure as code
  • Develop tools and utilities that reduce toil and empower operations and development teams to manage services independently
  • Apply AI-assisted and agentic workflows to development, support, and investigation work
  • Collaboration & Leadership
  • Collaborate with cross-functional development teams to support, enhance, and scale NaaS applications
  • Provide guidance and mentorship to junior engineers
  • Maintain clear documentation for processes and architecture

Requirements

  • Bachelor’s degree or equivalent in engineering, computer science, or related field.
  • 8+ years in software development, systems engineering, and/or networking
  • 5+ years of related experience required.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP), including compute, networking, and identity services
  • Strong automation and infrastructure-as-code skills: Terraform, Ansible, and Python
  • Experience running containerized workloads on Kubernetes
  • Working knowledge of modern observability and monitoring tooling (e.g., Datadog, CloudWatch, Grafana, Prometheus), including building dashboards and defining alerts
  • Demonstrated experience with incident management and blameless postmortems
  • Comfort using AI-assisted development and agentic tools as part of daily engineering practice
  • Understanding of network technologies including Internet, Ethernet, IPVPN, Edge Compute, and Optical transport
  • Strong listening and communication skills; able to operate with autonomy while knowing when to escalate

Preferred Qualifications:

  • Multi-cloud experience across AWS, Azure, and GCP
  • Asynchronous programming concepts and distributed systems design
  • Zero-downtime deployment strategies
  • High availability and multi-region architectures
  • Source control and CI/CD practices at scale
  • Experience applying agentic workflows to operational support and investigation

Benefits & conditions

This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.

Location Based Pay Ranges

$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY

$111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI

$116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA

Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We’re able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.

About the company

Lumen is the trusted network for the AI-powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high-performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.

At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.

This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI-ready connectivity, join us today., Life at Lumen is human and connected, even in a fast moving, AI-focused organization. We set clear expectations and trust people to meet them. With real support and shared accountability, teams collaborate better, move faster, and deliver meaningful outcomes.

Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win.

To learn more about Life at Lumen and how we live the Lumen 8, please visit

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all