Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
Design and maintain highly available production systems. -Define and manage SLIs, SLOs, and error budgets. -Automate operational tasks and eliminate manual processes. -Develop monitoring, alerting, and observability solutions. -Improve system performance, capacity, and resilience. -Lead incident response and root cause analysis. -Implement disaster recovery and continuity strategies. -Partner with development teams to improve application reliability.
Requirements
Bachelor’s with 12+ years of infrastructure/cloud engineering experience (or commensurate experience) -5-10+ years of engineering experience, with a strong background in Linux and Windows systems -Expertise in Kubernetes and container platforms -Experience working with cloud infrastructure environments -Proficiency in scripting languages such as Python and Go -Hands-on knowledge of Terraform and automation tools -Familiarity with monitoring platforms and incident management practices -Experience designing and managing CI/CD pipelines
Preferred Skills and Experience: -Kubernetes certifications -AWS/Azure certifications -DevOps certifications -ITIL preferred
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…