> Markdown version of [/jobs/ext/2581306-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2581306-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Ethos Group Inc - **Location:** Irving, TX, United States - **Contract:** Permanent contract - **Skills:** Application Performance Management, Systems Engineering, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Linux, DevOps, Monitoring of Systems, Python (Programming Language), Windows PowerShell, Reliability Engineering, Runbook, Software Engineering, Scripting, Google Cloud, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Terraform - **Published:** August 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ec1750f89aec76c6 ## About the Role * Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent experience * Experience in Site Reliability Engineering, DevOps, Systems Engineering, or Cloud Infrastructure roles * Strong understanding of Linux and cloud technologies * Direct hands-on experience with AWS, Azure, or Google Cloud Platform * Experience writing and maintaining infrastructure-as-code tools such as Azure resource templates, Terraform, or CloudFormation * Direct hands-on experience supporting CI/CD pipelines and automation tools * Direct hands-on experience with production grade Kubernetes deployments * Experience with monitoring and observability platforms * Strong scripting skills in Python, PowerShell, Bash, or similar languages * Excellent troubleshooting and problem-solving abilities * Strong communication and collaboration skills Preferred Qualifications * Direct hands-on experience supporting large-scale production environments * Deep knowledge of networking, security, and cloud architecture principles * Experience with incident management and root cause analysis * Relevant cloud, Kubernetes, or DevOps certifications ## Description Ethos Group is seeking a talented and proactive Site Reliability Engineer (SRE) to join our growing technology team. This role is ideal for an engineer who is passionate about building highly available, scalable, and reliable systems while partnering closely with development and infrastructure teams to improve platform performance, automation, and operational excellence. As a Site Reliability Engineer, you will play a critical role in maintaining system uptime, optimizing application performance, automating operational processes, and supporting a modern cloud-based technology environment. Key Responsibilities * Design, implement, and maintain reliable, scalable, and secure cloud infrastructure * Monitor application and system performance to ensure maximum availability * Automate operational tasks and deployment processes * Respond to incidents, troubleshoot production issues, and drive root cause analysis * Develop and maintain monitoring, alerting, and observability solutions * Partner with software engineering teams to improve application reliability and performance * Support CI/CD pipelines and deployment automation initiatives * Create operational documentation, runbooks, and best practices * Participate in on-call rotations and incident response activities * Continuously identify opportunities for system improvements and increased efficiency ## Related Videos - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)