> Markdown version of [/jobs/ext/2393578-engineer-site-reliability](https://www.wearedevelopers.com/jobs/ext/2393578-engineer-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineer, Site Reliability - **Company:** T-Mobile Us, Inc. - **Location:** Frisco, TX, United States - **Contract:** Permanent contract - **Skills:** Cloud Computing, Continuous Integration, HP Systems Insight Manager, Performance Tuning, Reliability Engineering, Software Deployment, Software Engineering, Scripting, Cloud Platform System, Reliability of Systems, Information Technology - **Published:** August 5, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/engineer-site-reliability-frisco-tx-usa-58837029 ## About the Role performance * Participate in additional duties/projects as assigned by management Tasks * Bachelor in Computer Science (or Engineering) with 3 years of related experience; advanced degree with 1 year related experience; or equivalent combination * Experience in CI/CD pipelines for software deployment (preferred) * Experience with cloud-native platforms and solutions (preferred) * Mentoring or guiding teams in reliability engineering practices (preferred) * Knowledge areas: Application Monitoring, Automation, CI/CD, Capacity Planning, Cloud Computing, Incident Management, Performance Tuning, Scripting, System Reliability Key requirements * medical, dental and vision insurance * 401(k) * paid time off and holidays * employee stock grants and ESPP * tuition assistance * mobile service and home internet discounts ## Description Experteer Overview In this role you will bolster the reliability and resilience of digital infrastructure by automating tasks, monitoring health, and steering incident responses. You'll work with cross-functional teams to drive uptime and reduce manual interventions, shaping efficient deployment and operational workflows. The position centers on automation, cloud-native practices, and proactive incident management to sustain high service quality. This is an opportunity to influence platform robustness at scale and contribute to a culture of reliability. Compensation / Benefits * Automate processes to improve system reliability and cut manual tasks * Monitor systems proactively to prevent incidents and ensure continuity * Streamline software development and deployment processes for operational efficiency * Develop scripts and tools to reduce routine workloads * Manage incident response for rapid recovery and minimal disruption * Adopt new technologies to sustain system robustness and performance * Participate in additional duties/projects as assigned by management Tasks * Bachelor in Computer Science (or Engineering) with 3 years of related experience; advanced degree with 1 year related experience; or equivalent combination * Experience in CI/CD pipelines for software deployment (preferred) * Experience with cloud-native platforms and solutions (preferred) * Mentoring or guiding teams in reliability engineering practices (preferred) * Knowledge areas: Application Monitoring, Automation, CI/CD, Capacity Planning, Cloud Computing, Incident Management, Performance Tuning, Scripting, System Reliability Key requirements * medical, dental and vision insurance * 401(k) * paid time off and holidays * employee stock grants and ESPP * tuition assistance * mobile service and home internet discounts ## Related Videos - [Unlocking the potential of Digital & IT at Vodafone](https://www.wearedevelopers.com/videos/602-unlocking-the-potential-of-digital-it-at-vodafone) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Intermediate Bitcoin Script](https://www.wearedevelopers.com/videos/25-intermediate-bitcoin-script) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)