> Markdown version of [/jobs/ext/3449879-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3449879-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Eliassen Group - **Location:** Westlake, TX, United States - **Salary:** $124,800.0 - $135,200.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Data Analysis, Application Portfolio Management, Systems Engineering, Microsoft Azure, Cloud Engineering, Databases, Continuous Integration, Distributed Systems, Groovy, Identity and Access Management, Python (Programming Language), Node.Js, Systems Development Life Cycle, Reliability Engineering, Ansible, Prometheus, Shell Script, Software Engineering, Datadog, Data Logging, Cloud Platform System, Grafana, Reliability of Systems, Kubernetes, Infrastructure Automation Frameworks, Cloud Migration, Terraform, Splunk, Jenkins - **Published:** September 30, 2026 - **Apply:** https://www.juju.com/job/15_a50f3c8d3 ## About the Role Due to client requirements, applicants must be willing and able to work on a w2 basis. For our w2 consultants, we offer a great benefits package that includes Medical, Dental, and Vision benefits, 401k with company matching, and life insurance., * 5+ years of experience supporting or building large-scale, multi-tiered distributed systems * 1-2 years of cloud development or cloud migration experience * 2-4 years of software development experience with a focus on automation and SDLC practices * Prior on-call experience supporting production systems and running incidents * Strong SRE, systems engineering, or software engineering background * Hands-on experience with AWS cloud environments (EKS preferred); Azure (AKS) experience is a plus * Solid Kubernetes experience in production environments * Experience with infrastructure as code tools such as Terraform, Ansible, Chef, IAM, or ARM * Proficiency in automation and scripting (Python, Shell scripting, Node.js, JavaScript, or Java) * CI/CD pipeline experience using tools such as Jenkins and Groovy * Proven experience supporting distributed, highly concurrent, service-based architectures * Hands-on experience with observability platforms such as Datadog, Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, or Splunk * Experience managing and analyzing large datasets, metrics, and logs to improve system reliability * Demonstrated ability to maintain scalability and resiliency in complex environments What Sets You Apart * Strong systems-thinking mindset with a passion for reliability and automation * Comfortable operating in high-pressure production environments * Excellent communication skills with the ability to engage both technical and non-technical partners * Proven ability to learn new tools and practices and introduce them effectively to engineering teams * Collaborative approach to working with diverse teams across locations and time zones, Bachelor's degree in a technology-related field or equivalent experience. Master's degree is a plus. ## Description We are seeking a Site Reliability Engineer to join a large-scale enterprise infrastructure organization within the financial services industry. This team is responsible for ensuring the reliability, scalability, and resilience of thousands of production systems supporting critical business functions. This role blends systems engineering, software development, and operations excellence, with a strong emphasis on automation, infrastructure as code, observability, and incident response. You will work in a highly collaborative SRE environment focused on building reliability into platforms rather than reacting to failures and work towards driving operations excellence, automation, and resiliency at scale. The role focuses on building reliable services through infrastructure as code, observability, and chaos testing while guiding developers and improving production insights. The Production Services team supports a broad application portfolio with follow-the-sun on-call rotation and delivers platform, application, batch, cloud, UI, middle tier, database, mainframe, release, and performance engineering services., * Design, build, and operate highly available, resilient infrastructure at enterprise scale * Drive automation-first solutions across operations, incident management, and environment management * Support production systems through a roster-based on-call rotation (follow-the-sun model) * Lead and participate in incident response, triage, and root cause analysis under pressure * Implement and enhance observability solutions including monitoring, logging, metrics, and alerting * Build and maintain infrastructure as code for cloud and platform services * Partner closely with application and platform teams to provide production insights and developer guidance * Continuously improve system reliability using resiliency engineering, chaos testing, and performance engineering practices ## Related Videos - [GitLab CI pipelines for a whole company](https://www.wearedevelopers.com/videos/143-gitlab-ci-pipelines-for-a-whole-company) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Give your build some love, it will give it back!](https://www.wearedevelopers.com/videos/514-give-your-build-some-love-it-will-give-it-back) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)