> Markdown version of [/jobs/ext/1449709-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1449709-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Fmr LLC - **Location:** Westlake, TX, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Amazon Web Services, Data Analysis, Build Automation, Microsoft Azure, Cloud Computing, Continuous Availability, Continuous Integration, Linux, Programming Tools, Fault Tolerance, Monitoring of Systems, Information Technology Operations, Windows PowerShell, Software Architecture, Reliability Engineering, Alwayson, Datadog, Scripting, Load Balancing, Grafana, Kubernetes, Software Coding, Splunk, Jenkins - **Published:** July 26, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f89430a3dff79b25 ## About the Role * 2 plus years of experience in systems and platform operations and technology management * Experience in Cloud computing(Azure and AWS), VMs, Windows and Linux. * Experience in managing Kubernetes cluster administration and expected to have good experience troubleshooting Kubernetes. * Build automation, scripts, and lightweight tools to eliminate repetitive manual work and improve operational efficiency. * Experience in Python scripting and PowerShell skills are highly preferred. * Experience with analytics and monitoring tools such as Grafana, Splunk and Datadog * Experience supporting 24/7, continuous availability production and managed environments. * Good understanding of software architecture helps empower software developers and engineers to build platforms with greater resiliency and fault tolerance in mind. * SRE (Site Reliability Engineering) principles a plus - resiliency, observability, and governance gating * Participate in on-call rotations, respond to incidents, execute runbooks, and ensure clear communication and handoffs. * Experience with continuous integration tools, such as Jenkins and AWX * Good to have knowledge of networking, firewalls and load balancers. * Knowledge of best practices for IT operations in an always-on, always-available service model * Ability to work across teams to continuously analyze system performance in production, troubleshoot consumer reported issues, and proactively identify areas in need of optimization. * An ability to lead change in a creative and collaborative manner with business and technical partners. * Great communication, collaboration, and interpersonal skills * Optional certifications: AWS, Azure related credentials., Ideal candidates will have a background in Site Reliability with a strong desire to expand into the other domain, or prior experience as a Site Reliability Engineer. We are looking for a systems-thinking Site Reliability Engineer who has helped teams scale through production insights, operational automation, developer enablement, real-time metrics, and continuous improvement. ## Description We are looking for Site Reliability Engineer who solve operational problems by building software. In this role, you will improve reliability, reduce toil, and enhance production systems by writing code, building automations, and leveraging modern AI-assisted development tools. The position requires a deep understanding of application and infrastructure support and expert with supporting cloud computing environments. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)