Senior Site Reliability Engineer

Castleton Commodities International LLC
Houston, TX, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Automation of Tests Bash Shell Computer Programming Continuous Integration Linux DevOps Github Python (Programming Language) Nagios Reliability Engineering Prometheus
+16 more
Runbook Datadog Data Logging Scripting Computer Network Technologies Autoscaling Delivery Pipeline Grafana Mttr Cloudformation Containerization Gitlab-ci Kubernetes Terraform Jenkins Golang

Job description

Experteer Overview In this role you will strengthen the reliability and scalability of critical cloud-native infrastructure. You’ll partner with Engineering, Security, and Infrastructure teams to design resilient architectures, drive IaC and CI/CD maturity, and measure reliability through SLOs/SLIs. You will lead DR and BCP initiatives, validate recovery objectives, and run structured testing to ensure readiness. This position offers the chance to shape operations for high-availability services and contribute to a culture of continuous improvement. Compensation / Benefits * Own and improve service reliability via SLO/SLI, error budgets, and best practices * Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR * Lead incident response through on-call improvements, runbooks, RCA, and preventative actions * Collaborate with application teams to enhance performance, capacity planning, and resiliency under failure * Design and operate highly available, multi-AZ/multi-region cloud architectures * Implement resilient patterns for compute, storage, networking, and managed services * Drive cloud governance practices with security and platform teams * Build and maintain IaC modules (Terraform, CloudFormation, CDK) for auditable infrastructure * Develop and optimize CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline) * Promote DevOps: versioned infra, automated testing, immutable deployments, progressive delivery * Ensure environment consistency across dev/test/stage/prod and drift remediation * Collaborate on defining RTO/RPO and design DR architectures and procedures * Coordinate and execute structured DR tests and document outcomes * Maintain DR runbooks, dependency maps, and recovery checklists; drive gap remediation * Produce metrics and reporting on DR readiness and continuous improvement actions Tasks * 7+ years in SRE, DevOps, or related roles * Strong observability experience (Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft) * Hands-on AWS experience in production environments * Proficiency with IaC (Terraform and/or CloudFormation/CDK) * CI/CD and automation expertise (pipeline design, testing automation) * Experience with RTO/RPO definition and BCP/DR testing * Kubernetes and auto-scaling container platforms (EKS, ECS) * Strong Linux fundamentals and networking knowledge * Scripting/programming (Python, Go, Bash) * Ability to write clear runbooks and post-incident reports * Ability to work in fast-paced environments and beyond-hours availability Key requirements * Competitive medical, dental, retirement and life insurance * Employee assistance & wellness programs * Parental and family leave policies * Charity volunteer days and matching program * Tuition assistance & reimbursement * Quarterly Innovation & Collaboration Awards

Requirements

available, multi-AZ/multi-region cloud architectures * Implement resilient patterns for compute, storage, networking, and managed services * Drive cloud governance practices with security and platform teams * Build and maintain IaC modules (Terraform, CloudFormation, CDK) for auditable infrastructure * Develop and optimize CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline) * Promote DevOps: versioned infra, automated testing, immutable deployments, progressive delivery * Ensure environment consistency across dev/test/stage/prod and drift remediation * Collaborate on defining RTO/RPO and design DR architectures and procedures * Coordinate and execute structured DR tests and document outcomes * Maintain DR runbooks, dependency maps, and recovery checklists; drive gap remediation * Produce metrics and reporting on DR readiness and continuous improvement actions Tasks * 7+ years in SRE, DevOps, or related roles * Strong observability experience (Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft) * Hands-on AWS experience in production environments * Proficiency with IaC (Terraform and/or CloudFormation/CDK) * CI/CD and automation expertise (pipeline design, testing automation) * Experience with RTO/RPO definition and BCP/DR testing * Kubernetes and auto-scaling container platforms (EKS, ECS) * Strong Linux fundamentals and networking knowledge * Scripting/programming (Python, Go, Bash) * Ability to write clear runbooks and post-incident reports * Ability to work in fast-paced environments and beyond-hours availability Key requirements * Competitive medical, dental, retirement and life insurance * Employee assistance & wellness programs * Parental and family leave policies * Charity volunteer days and matching program * Tuition assistance & reimbursement * Quarterly Innovation & Collaboration Awards

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all