Lead Engineer - Site Reliability

Frontier Airlines
Denver, CO, United States
28 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$110,114.0 - $146,157.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Computing Platforms Software as a Service Cloud Computing Cloud Engineering Configuration Management Continuous Integration Disaster Recovery Distributed Systems PCI Data Security Standards Reliability Engineering Software Engineering
+16 more
Web Platforms Data Logging System Availability Generative AI Infrastructure as Code (IaC) Cloudformation Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Cloud Migration ArcSight Event Correlation Terraform Dynatrace Devsecops Microservices

Job description

As a Lead Site Reliability Engineer (SRE), you will serve as a senior technical leader responsible for ensuring the reliability, scalability, performance, and operational excellence of mission-critical technology platforms. You will partner with engineering, platform, security, and operations teams to build resilient systems, drive automation, improve service availability, and establish reliability-focused engineering practices across the organization.

This role requires deep expertise in AWS, Kubernetes, observability, distributed systems, DevSecOps, and Site Reliability Engineering principles, with the ability to influence engineering teams and operational strategies across the enterprise.

Essential Functions

  • Lead the design and implementation of highly available, resilient, and scalable cloud platforms leveraging AWS services and cloud-native technologies.
  • Define and execute reliability engineering strategies, service level objectives (SLOs), service level indicators (SLIs), and error budgets across critical business services.
  • Architect and operate Kubernetes-based platforms supporting containerized and microservices-based applications.
  • Establish enterprise standards for reliability, availability, disaster recovery, resiliency testing, and operational excellence.
  • Partner with software engineering, platform engineering, security, and operations teams to improve service reliability and reduce operational risk.
  • Drive adoption of DevSecOps practices, CI/CD automation, Infrastructure as Code (IaC), and self-service platform capabilities.
  • Define and implement observability standards including monitoring, logging, distributed tracing, synthetic monitoring, and operational intelligence.
  • Lead incident response, post-incident reviews, root cause analyses (RCA), and continuous improvement initiatives.
  • Identify and eliminate single points of failure through proactive reliability engineering and resilient architecture practices.
  • Design and implement automated remediation, self-healing capabilities, and proactive alerting solutions to improve operational efficiency.
  • Lead capacity planning, performance engineering, availability management, and scalability assessments.
  • Partner with FinOps and engineering teams to optimize cloud resource utilization while maintaining reliability objectives.
  • Evaluate emerging technologies, AIOps, Generative AI, and intelligent automation solutions to improve platform reliability and operational effectiveness.
  • Mentor SREs, platform engineers, and software engineers in reliability practices, observability, automation, and operational excellence.
  • Develop executive-level reliability roadmaps, operational strategies, and platform investment recommendations.

Additional Responsibilities

  • Act as a trusted advisor to executive leadership on reliability strategy, operational risk management, and service resiliency initiatives.
  • Support platform evaluations, architecture reviews, and technology decisions through the lens of reliability, scalability, and maintainability.
  • Provide technical leadership during major incidents, business-critical outages, and high-severity escalations.
  • Assist with compliance initiatives including PCI-DSS, SOC 2, security governance, and operational resilience controls.
  • Contribute to enterprise-wide cloud transformation, platform engineering, and application modernization programs.
  • Champion a culture of reliability, automation, operational ownership, and continuous improvement across engineering teams., General Direction: The incumbent normally receives little instruction on day-to-day work and receives general instructions on new assignments.

Requirements

  • Bachelor’s degree in computer science, Engineering, Information Technology, or a related discipline.
  • 10 years of experience in software engineering, cloud architecture, infrastructure engineering, or enterprise architecture.
  • 5 years of hands-on AWS architecture and cloud transformation experience.
  • Proven success leading large-scale cloud migrations and modernization initiatives.
  • Experience designing and supporting highly available, mission-critical, customer-facing platforms.
  • Deep understanding of Kubernetes, container orchestration, microservices, and distributed systems.
  • Extensive experience with DevSecOps, CI/CD pipelines, Infrastructure as Code, and automation frameworks.
  • Strong knowledge of Site Reliability Engineering (SRE), operational excellence, and platform reliability practices.
  • Experience implementing cloud governance, FinOps, and cost optimization programs.

Preferred Qualifications

  • Experience supporting large-scale enterprise, eCommerce, aviation, travel, SaaS, or high-volume digital platforms.
  • Experience building internal developer platforms and platform engineering capabilities.
  • AWS Professional and/or Kubernetes certifications.
  • Experience with AIOps, intelligent automation, and reliability analytics.

Knowledge, Skills and Abilities

  • Reliability & Platform Technologies

  • Amazon Web Services (AWS)
  • Kubernetes (EKS)
  • Container Platforms
  • Platform Engineering
  • Infrastructure as Code (Terraform, CloudFormation)
  • Cloud Networking and Security
  • High Availability & Disaster Recovery
  • Performance Engineering & Capacity Planning

Site Reliability Engineering

  • Service Level Indicators (SLIs)
  • Service Level Objectives (SLOs)
  • Error Budget Management
  • Reliability Engineering
  • Availability & Resiliency Design
  • Chaos Engineering
  • Incident Response & Disaster Recovery
  • Root Cause Analysis (RCA)

  • Observability & Operations

  • Monitoring & Alerting
  • Distributed Tracing
  • Centralized Logging
  • Synthetic Monitoring
  • Operational Intelligence
  • AIOps & Event Correlation
  • Incident & Problem Management
  • Change & Release Management

  • DevSecOps & Automation

  • CI/CD Automation
  • DevSecOps Practices
  • Infrastructure Automation
  • Platform Automation
  • Self-Healing Systems
  • Configuration Management
  • Reliability Automation

Leadership & Business Skills

  • Operational Strategy & Reliability Roadmaps
  • Cloud Financial Management (FinOps)
  • Executive Communication
  • Risk Management & Mitigation
  • Technical Leadership Without Direct Authority
  • Cross-Functional Collaboration
  • Major Incident Leadership
  • Continuous Improvement Leadership

Equipment Operated

Standard office equipment, including PC, copier, fax machine, printer

Benefits & conditions

Depending on role and eligibility, available benefits and programs may include:

  • Medical, dental and vision coverage
  • 401(k) retirement savings options
  • Paid holidays, vacation time and sick time
  • Travel privileges on Frontier Airlines and participating partner airlines, based on current program rules and availability
  • Buddy passes, based on eligibility and program rules
  • Travel-related discounts and employee discounts on select products, services and vendors
  • A hybrid schedule for eligible headquarters roles based in Denver, Colorado
  • Business casual dress options for eligible corporate and support roles
  • Employee support programs and resources, including the HOPE League, Frontier Airlines’ nonprofit organization

About Frontier Airlines

Frontier Airlines is a Denver-based airline serving destinations across the United States and select international markets. Our people support every part of the travel journey, from airport operations and flight crews to aircraft maintenance, customer support, corporate teams and more.

We’re focused on delivering meaningful value by making travel more accessible, practical and easy to personalize for our customers.

About the company

At Frontier, our mission is to Make Every Flight Count. That mission guides how we support our customers, our people and the operation every day.

As a Frontier employee, your work connects to more than a single role. Whether you’re supporting flights, helping customers, maintaining aircraft, leading teams or working behind the scenes, you help create a travel experience that is safe, reliable and built around value.

Our work is guided by our core values: customer first, safety always, operational excellence and one team. These values shape how we make decisions, support each other and deliver for the people who count on us.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:32 min

Overview of Terraform and Terraform Cloud features

Devlin Duldulao · LIVE

Videos

See all

Related articles

See all