Site Reliability Engineer

Insight Global
Arlington, VA, United States
15 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Microsoft Windows Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Linux DevOps Disaster Recovery Python (Programming Language) Reliability Engineering Scripting Delivery Pipeline
+3 more
Containerization Kubernetes Terraform

Job description

Design and maintain highly available production systems. -Define and manage SLIs, SLOs, and error budgets. -Automate operational tasks and eliminate manual processes. -Develop monitoring, alerting, and observability solutions. -Improve system performance, capacity, and resilience. -Lead incident response and root cause analysis. -Implement disaster recovery and continuity strategies. -Partner with development teams to improve application reliability.

Requirements

Bachelor’s with 12+ years of infrastructure/cloud engineering experience (or commensurate experience) -5-10+ years of engineering experience, with a strong background in Linux and Windows systems -Expertise in Kubernetes and container platforms -Experience working with cloud infrastructure environments -Proficiency in scripting languages such as Python and Go -Hands-on knowledge of Terraform and automation tools -Familiarity with monitoring platforms and incident management practices -Experience designing and managing CI/CD pipelines

Preferred Skills and Experience: -Kubernetes certifications -AWS/Azure certifications -DevOps certifications -ITIL preferred

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Loading talks and stories from around this role…