Lead Site Reliability Engineer

Lumen Inc
Rapid City, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 141K

Job location

Remote
Rapid City, United States of America

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Application Layers
Application Lifecycle Management
Bash
Cloud Computing
Code Generation
Databases
Noise Reduction
Software Design Patterns
DevOps
Fault Tolerance
Github
Monitoring of Systems
Python
Pattern Recognition
Performance Tuning
Reliability Engineering
Software Engineering
Datadog
Data Logging
Scripting (Bash/Python/Go/Ruby)
Spring-boot
Containerization
Gitlab-ci
Kubernetes
Infrastructure Automation Frameworks
Low Latency
Cloudwatch
Terraform
Docker
Jenkins
Microservices

Job description

We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering practices.

The Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.

This role will collaborate with key stakeholders across the engineering organization - including product owners, developers, and testers - to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization., Production Support & Incident Management

  • Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
  • Provide Tier 3 support for issues across portal services by troubleshooting and resolving technical issues in test and production environments.
  • Lead root cause analysis and post-mortem processes, incorporating AI-assisted analysis and pattern detection to ensure continuous improvement.

Performance Optimization

  • Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
  • Implement tuning strategies across application layers, databases, and infrastructure.
  • Drive initiatives to improve latency, throughput, and resource utilization.

Monitoring & Observability

  • Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
  • Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
  • Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.

Automation & Infrastructure as Code

  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery.
  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
  • Use Terraform, or similar IaC tools to manage AWS resources.

Reliability Engineering

  • Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
  • Champion SRE principles such as SLIs, SLOs, and error budgets.
  • Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.

Collaboration & Communication

  • Work closely with software engineers, DevOps, and product teams to align reliability goals.
  • Document processes, runbooks, and best practices for knowledge sharing.
  • Provide mentorship and guidance on reliability and operational excellence.

Requirements

  • 5 years overall professional experience in SRE, DevOps, or infrastructure engineering roles.
  • Experience with Terraform, or similar IaC tools to manage Cloud resources.
  • Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks.
  • Experience with CI/CD pipelines and tools like GitHub Actions, Jenkins or GitLab CI.
  • Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, Datadog).
  • Familiarity with containerization and orchestration (Docker, Kubernetes).
  • Excellent AI and problem-solving skills, and a proactive mindset.

Preferred Qualifications:

  • Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.).
  • Certifications in AWS or related technologies are a plus.
  • Experience of application development using Java Microservices and Spring Boot framework
  • Experience with Agile/SCRUM Methodologies and development practices

Benefits & conditions

This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.

Location Based Pay Ranges

$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY $111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI $116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA

Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.

Learn more about Lumen's:

  • Benefits
  • Bonus Structure

About the company

Lumen is the trusted network for the AI-powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high-performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities. At Lumen, you'll work on infrastructure customers rely on today and build for what's next, where performance, security, and resilience matter. This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you're ready to take ownership, deliver meaningful impact, and help shape the future of AI-ready connectivity, join us today., Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win., Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.

Apply for this position