> Markdown version of [/jobs/ext/2564957-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2564957-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** Selby Jennings - **Location:** Wilmington, NC, United States - **Experience:** Expert - **Salary:** $175,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Configuration Management, Computer Programming, Continuous Integration, Disaster Recovery, Monitoring of Systems, Reliability Engineering, Software Engineering, Workflow Management Systems, Data Logging, Git, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Terraform, Software Version Control, Docker - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e0c3074ed60a7215 ## About the Role * Bachelor's degree in Computer Science or a related field, or equivalent professional experience. * Advanced degree in Computer Science or a related discipline is preferred., * AWS cloud infrastructure and services * Kubernetes and container orchestration platforms * Infrastructure as Code (Terraform or similar tools) * CI/CD and deployment automation * Git and modern version control practices * Containerization technologies (Docker) * Monitoring, logging, and observability platforms * Scripting and programming experience * Database administration and management * Incident and problem management * Security, compliance, and cloud governance Preferred Experience * Experience with enterprise monitoring and observability platforms * Cloud networking and security best practices * Disaster recovery and resilience planning * Workflow automation and orchestration tools * Experience in insurance, financial services, or other regulated industries * AWS certifications or equivalent cloud certifications preferred ## Description The company is seeking a highly skilled and motivated Lead Site Reliability Engineer to play a key role in designing, implementing, and maintaining reliable, scalable, and high-performing cloud infrastructure within AWS. This individual will work closely with software engineering, operations, and cross-functional teams to improve platform reliability, enhance developer productivity, and drive operational excellence through automation, monitoring, and incident response practices., * Lead and mentor a team of Site Reliability Engineers, including both full-time employees and contractors. * Prioritize, assign, and review technical work while providing guidance and feedback on code and infrastructure changes. * Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in AWS. * Build and support monitoring, alerting, and observability solutions to ensure platform health and uptime. * Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools. * Develop and enhance CI/CD pipelines to improve deployment efficiency and software delivery. * Lead incident response efforts, conduct root cause analysis, and implement long-term solutions. * Partner with engineering teams to optimize performance, reliability, scalability, and cloud costs. * Promote operational best practices across infrastructure and application environments. * Develop and maintain disaster recovery and business continuity capabilities. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)