Lead Site Reliability Engineer

Selby Jennings
Wilmington, DE, United States
5 days ago
Apply on www.efinancialcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Cloud Computing Configuration Management Continuous Integration Disaster Recovery Monitoring of Systems Uptime Reliability Engineering Workflow Management Systems Data Logging Git Containerization
+7 more
Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Terraform Software Version Control Docker

Job description

response practices. Key Responsibilities Lead and mentor a team of Site Reliability Engineers, including both full-time employees and contractors. Prioritize, assign, and review technical work while providing guidance and feedback on code and infrastructure changes. Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in AWS. Build and support monitoring, alerting, and observability solutions to ensure platform health and uptime. Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools. Develop and enhance CI/CD pipelines to improve deployment efficiency and software delivery. Lead incident response efforts, conduct root cause analysis, and implement long-term solutions. Partner with engineering teams to optimize performance, reliability, scalability, and cloud costs. Promote operational best practices across infrastructure and application environments. Develop and maintain disaster recovery and business

Requirements

continuity capabilities. Qualifications Bachelor’s degree in Computer Science or a related field, or equivalent professional experience. Advanced degree in Computer Science or a related discipline is preferred. Required Technical Skills AWS cloud infrastructure and services Kubernetes and container orchestration platforms Infrastructure as Code (Terraform or similar tools) CI/CD and deployment automation Git and modern version control practices Containerization technologies (Docker) Monitoring, logging, and observability platforms Scripting and programming experience Database administration and management Incident and problem management Security, compliance, and cloud governance Preferred Experience Experience with enterprise monitoring and observability platforms Cloud networking and security best practices Disaster recovery and resilience planning Workflow automation and orchestration tools Experience in insurance, financial services, or other regulated industries AWS certifications or equivalent cloud certifications preferred

About the company

Lead Site Reliability Engineer About the Company This organization is focused on modernizing the mortgage insurance industry through a technology-first approach. Rather than relying on legacy systems and manual processes, the company leverages software, automation, artificial intelligence, analytics, and scalable operating models to deliver better customer experiences and drive business efficiency. Its mission is to build a more modern, data-driven insurance platform designed for long-term growth and innovation. About the Role The company is seeking a highly skilled and motivated Lead Site Reliability Engineer to play a key role in designing, implementing, and maintaining reliable, scalable, and high-performing cloud infrastructure within AWS. This individual will work closely with software engineering, operations, and cross-functional teams to improve platform reliability, enhance developer productivity, and drive operational excellence through automation, monitoring, and incident

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.efinancialcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:22 min

Leveraging unique cultural backgrounds in engineering design

Ixchel Ruiz · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all