Senior Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
Experteer Overview In this role you will strengthen the reliability and scalability of critical cloud-native infrastructure. You’ll partner with Engineering, Security, and Infrastructure teams to design resilient architectures, drive IaC and CI/CD maturity, and measure reliability through SLOs/SLIs. You will lead DR and BCP initiatives, validate recovery objectives, and run structured testing to ensure readiness. This position offers the chance to shape operations for high-availability services and contribute to a culture of continuous improvement. Compensation / Benefits * Own and improve service reliability via SLO/SLI, error budgets, and best practices * Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR * Lead incident response through on-call improvements, runbooks, RCA, and preventative actions * Collaborate with application teams to enhance performance, capacity planning, and resiliency under failure * Design and operate highly available, multi-AZ/multi-region cloud architectures * Implement resilient patterns for compute, storage, networking, and managed services * Drive cloud governance practices with security and platform teams * Build and maintain IaC modules (Terraform, CloudFormation, CDK) for auditable infrastructure * Develop and optimize CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline) * Promote DevOps: versioned infra, automated testing, immutable deployments, progressive delivery * Ensure environment consistency across dev/test/stage/prod and drift remediation * Collaborate on defining RTO/RPO and design DR architectures and procedures * Coordinate and execute structured DR tests and document outcomes * Maintain DR runbooks, dependency maps, and recovery checklists; drive gap remediation * Produce metrics and reporting on DR readiness and continuous improvement actions Tasks * 7+ years in SRE, DevOps, or related roles * Strong observability experience (Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft) * Hands-on AWS experience in production environments * Proficiency with IaC (Terraform and/or CloudFormation/CDK) * CI/CD and automation expertise (pipeline design, testing automation) * Experience with RTO/RPO definition and BCP/DR testing * Kubernetes and auto-scaling container platforms (EKS, ECS) * Strong Linux fundamentals and networking knowledge * Scripting/programming (Python, Go, Bash) * Ability to write clear runbooks and post-incident reports * Ability to work in fast-paced environments and beyond-hours availability Key requirements * Competitive medical, dental, retirement and life insurance * Employee assistance & wellness programs * Parental and family leave policies * Charity volunteer days and matching program * Tuition assistance & reimbursement * Quarterly Innovation & Collaboration Awards
Requirements
available, multi-AZ/multi-region cloud architectures * Implement resilient patterns for compute, storage, networking, and managed services * Drive cloud governance practices with security and platform teams * Build and maintain IaC modules (Terraform, CloudFormation, CDK) for auditable infrastructure * Develop and optimize CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline) * Promote DevOps: versioned infra, automated testing, immutable deployments, progressive delivery * Ensure environment consistency across dev/test/stage/prod and drift remediation * Collaborate on defining RTO/RPO and design DR architectures and procedures * Coordinate and execute structured DR tests and document outcomes * Maintain DR runbooks, dependency maps, and recovery checklists; drive gap remediation * Produce metrics and reporting on DR readiness and continuous improvement actions Tasks * 7+ years in SRE, DevOps, or related roles * Strong observability experience (Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft) * Hands-on AWS experience in production environments * Proficiency with IaC (Terraform and/or CloudFormation/CDK) * CI/CD and automation expertise (pipeline design, testing automation) * Experience with RTO/RPO definition and BCP/DR testing * Kubernetes and auto-scaling container platforms (EKS, ECS) * Strong Linux fundamentals and networking knowledge * Scripting/programming (Python, Go, Bash) * Ability to write clear runbooks and post-incident reports * Ability to work in fast-paced environments and beyond-hours availability Key requirements * Competitive medical, dental, retirement and life insurance * Employee assistance & wellness programs * Parental and family leave policies * Charity volunteer days and matching program * Tuition assistance & reimbursement * Quarterly Innovation & Collaboration Awards
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
DevOps Engineer Salary [2023]
Dev Digest 121 - AI goes offline