> Markdown version of [/jobs/ext/1969046-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1969046-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Castleton Commodities International LLC - **Location:** Houston, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Automation of Tests, Bash Shell, Computer Programming, Continuous Integration, Linux, DevOps, Github, Python (Programming Language), Nagios, Reliability Engineering, Prometheus, Runbook, Datadog, Data Logging, Scripting, Computer Network Technologies, Autoscaling, Delivery Pipeline, Grafana, Mttr, Cloudformation, Containerization, Gitlab-ci, Kubernetes, Terraform, Jenkins, Golang - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-site-reliability-engineer-houston-tx-usa-58837174 ## About the Role available, multi-AZ/multi-region cloud architectures * Implement resilient patterns for compute, storage, networking, and managed services * Drive cloud governance practices with security and platform teams * Build and maintain IaC modules (Terraform, CloudFormation, CDK) for auditable infrastructure * Develop and optimize CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline) * Promote DevOps: versioned infra, automated testing, immutable deployments, progressive delivery * Ensure environment consistency across dev/test/stage/prod and drift remediation * Collaborate on defining RTO/RPO and design DR architectures and procedures * Coordinate and execute structured DR tests and document outcomes * Maintain DR runbooks, dependency maps, and recovery checklists; drive gap remediation * Produce metrics and reporting on DR readiness and continuous improvement actions Tasks * 7+ years in SRE, DevOps, or related roles * Strong observability experience (Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft) * Hands-on AWS experience in production environments * Proficiency with IaC (Terraform and/or CloudFormation/CDK) * CI/CD and automation expertise (pipeline design, testing automation) * Experience with RTO/RPO definition and BCP/DR testing * Kubernetes and auto-scaling container platforms (EKS, ECS) * Strong Linux fundamentals and networking knowledge * Scripting/programming (Python, Go, Bash) * Ability to write clear runbooks and post-incident reports * Ability to work in fast-paced environments and beyond-hours availability Key requirements * Competitive medical, dental, retirement and life insurance * Employee assistance & wellness programs * Parental and family leave policies * Charity volunteer days and matching program * Tuition assistance & reimbursement * Quarterly Innovation & Collaboration Awards ## Description Experteer Overview In this role you will strengthen the reliability and scalability of critical cloud-native infrastructure. You'll partner with Engineering, Security, and Infrastructure teams to design resilient architectures, drive IaC and CI/CD maturity, and measure reliability through SLOs/SLIs. You will lead DR and BCP initiatives, validate recovery objectives, and run structured testing to ensure readiness. This position offers the chance to shape operations for high-availability services and contribute to a culture of continuous improvement. Compensation / Benefits * Own and improve service reliability via SLO/SLI, error budgets, and best practices * Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR * Lead incident response through on-call improvements, runbooks, RCA, and preventative actions * Collaborate with application teams to enhance performance, capacity planning, and resiliency under failure * Design and operate highly available, multi-AZ/multi-region cloud architectures * Implement resilient patterns for compute, storage, networking, and managed services * Drive cloud governance practices with security and platform teams * Build and maintain IaC modules (Terraform, CloudFormation, CDK) for auditable infrastructure * Develop and optimize CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline) * Promote DevOps: versioned infra, automated testing, immutable deployments, progressive delivery * Ensure environment consistency across dev/test/stage/prod and drift remediation * Collaborate on defining RTO/RPO and design DR architectures and procedures * Coordinate and execute structured DR tests and document outcomes * Maintain DR runbooks, dependency maps, and recovery checklists; drive gap remediation * Produce metrics and reporting on DR readiness and continuous improvement actions Tasks * 7+ years in SRE, DevOps, or related roles * Strong observability experience (Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft) * Hands-on AWS experience in production environments * Proficiency with IaC (Terraform and/or CloudFormation/CDK) * CI/CD and automation expertise (pipeline design, testing automation) * Experience with RTO/RPO definition and BCP/DR testing * Kubernetes and auto-scaling container platforms (EKS, ECS) * Strong Linux fundamentals and networking knowledge * Scripting/programming (Python, Go, Bash) * Ability to write clear runbooks and post-incident reports * Ability to work in fast-paced environments and beyond-hours availability Key requirements * Competitive medical, dental, retirement and life insurance * Employee assistance & wellness programs * Parental and family leave policies * Charity volunteer days and matching program * Tuition assistance & reimbursement * Quarterly Innovation & Collaboration Awards ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)