Lead Site Reliability Engineer

Liberty Personnel Services, Inc.
Philadelphia, PA, United States
3 months ago
Apply on libertyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Bash Shell Cloud Computing Continuous Integration DevOps Disaster Recovery Python (Programming Language) Reliability Engineering Datadog Cloud Platform System Git Flow Kubernetes Infrastructure Automation Frameworks
+1 more
Terraform

Job description

The Lead Site Reliability Engineer is a senior technical leader responsible for the reliability, availability, and operational excellence of a cloud-based infrastructure and distributed platform. This role owns uptime, SLAs, and incident response while driving long-term improvements in resilience, observability, and automation. The Lead SRE is hands-on and partners closely with platform, QA, and development teams.

This role suits an engineer who thrives in high-ownership environments, balancing real-time operations with strategic reliability initiatives. You’ll define operational standards, disaster recovery practices, and automation frameworks, while leading incidents and postmortems with clarity and accountability., * Own uptime, SLAs, and overall platform reliability

  • Lead incident response, root-cause analysis, and postmortems
  • Automate infrastructure, deployments, and operational workflows
  • Improve monitoring, alerting, and observability
  • Execute and evolve disaster recovery and business continuity plans
  • Optimize cloud and Kubernetes environments for scale and performance
  • Establish runbooks, operational standards, and reliability best practices
  • Provide technical leadership and mentorship

Requirements

  • 6+ years in SRE, DevOps, or Platform Engineering; 2+ years in a lead role
  • Strong experience supporting production systems with strict SLAs
  • Deep expertise in Kubernetes, containers, and cloud infrastructure
  • Proficiency with Terraform and modern IaC practices
  • Strong automation and scripting skills (Bash, Python, or Go)
  • Experience with CI/CD, GitOps, and observability tooling
  • Proven incident leadership and cross-functional communication skills

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on libertyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:53 min

Introduction to git flow and clean feature branches

Johannes Haux · World Congress 2022

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

Videos

See all

Related articles

See all