> Markdown version of [/jobs/ext/2976513-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2976513-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Inspyr Solutions - **Location:** Houston, TX, United States (Remote available) - **Experience:** Experienced - **Contract:** Temporary to permanent - **Skills:** Java (Programming Language), .NET Framework, Microsoft Windows, Application Programming Interfaces (APIs), Systems Engineering, Microsoft Azure, Bash Shell, Software as a Service, Continuous Integration, Linux, DevOps, Disaster Recovery, Distributed Systems, Python (Programming Language), Object-Oriented Software Development, Windows PowerShell, Reliability Engineering, Site Reliability Engineering Practices, Cloud Services, Ansible, Software Engineering, Datadog, System Availability, Kubernetes, Infrastructure Automation Frameworks, Hardware Infrastructure, Terraform, Microservices - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/0076a644-8615-4e96-bccb-2578a6bbc1a7 ## About the Role * 4+ years of experience in SRE, DevOps, Systems Engineering, or related roles * Strong Linux and Windows systems administration and troubleshooting skills * Hands-on experience with automation and scripting * Experience designing and operating monitoring, alerting, and observability solutions * Practical experience working in Azure environments * Strong analytical skills and a bias toward eliminating root causes, not symptoms * Ability to collaborate across application, infrastructure, and operations teams * Exposure to Kubernetes, microservices, or container orchestration * Hands-on experience with infrastructure as code tools such as Terraform or Ansible * Understanding of distributed systems and high availability design * Experience with SRE practices such as SLO based operations, runbook automation, or chaos testing ## Description This role exists to move the organization from reactive operations to engineered reliability. You will study how our most critical systems fail, particularly our internal applications and facility automation interfaces and design controls, automation, and observability that reduce incidents over time. Success in this role means fewer false alerts, faster recovery, less manual intervention, and systems that heal themselves when possible. You will work closely with application, infrastructure, and operations teams and participate directly in on call and incident response. What You Will Own * Definition and implementation of SLIs and SLOs that measure meaningful system health, not just availability * Observability across the full stack, correlating cloud services, APIs, and on premise facility operations * Automation to eliminate operational toil, including patching, data corrections, restarts, and recovery tasks * Development of self healing behaviors for common failure modes * Participation in on call rotations and leadership of blameless post incident reviews * Design and execution of disaster recovery tests across SaaS, cloud, and on premise environments This is hands on reliability engineering. The systems you improve will directly impact daily warehouse operations. Technical Environment * Hybrid environments spanning cloud and on-premise infrastructure * Azure cloud services * Software Development/OOP skills within either .NET/Java/Python * Observability tooling across logs, metrics, and alerting * Automation using Python, PowerShell, Bash, or Ansible * CI/CD tools and modern deployment practices * Exposure to containerized and distributed systems environments ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)