> Markdown version of [/jobs/ext/1841330-site-reliability-engineer-iii](https://www.wearedevelopers.com/jobs/ext/1841330-site-reliability-engineer-iii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer III - **Company:** Everforth Apex - **Location:** Plano, TX, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Java (Programming Language), Apache HTTP Server, Apache Tomcat, Applications Architecture, Application Performance Management, JIRA, Backup Devices, Unix, Command-Line Interface, Continuous Integration, DevOps, Middleware, Perl (Programming Language), Fault Tolerance, Apache Hadoop, Python (Programming Language), Microsoft SQL Server, Windows Servers, Oracle (Applications), Scrum Methodology, Red Hat Enterprise Linux, Ansible, Prometheus, Server Administration, Software Engineering, Scripting, ReactJS, System Availability, Grafana, Tanium Platform Expertise, AngularJS, Information Technology, Bitbucket, Front End Software Development, Restful APIs, Terraform, Splunk, Appdynamics, Dynatrace, Jenkins, Servicenow, Artifactory, Web Api - **Published:** July 21, 2026 - **Apply:** https://www.dice.com/job-detail/cd052bb0-db31-4a26-8f87-691eefcce465 ## About the Role Experience: 8 to 10 years of information technology experience, with over 6 years on a DevOps, SRE, or performance engineering team. Technical Skills: * Experience triaging production issues using APM tools (e.g., Dynatrace, AppDynamics, Prometheus, Grafana) and log aggregation tools (e.g., Splunk, ELK). * Strong experience in Java and front-end development (React JS, Angular). * Proficiency with Apache/Tomcat middleware and Java/RESTful services frameworks. * Backend database experience with Oracle, SQL Server, or Hadoop. * Strong scripting skills in Python, UNIX, or Perl/Shell. * Experience with CI/CD tools such as Bitbucket, JFrog Artifactory, Jenkins, Terraform, and Ansible. * Knowledge of SRE concepts like SLIs/SLOs and error budgets. * Experience with the Agile/Scrum methodology. Education: A college degree or equivalent work experience is required. Preferred Qualifications * Proficiency in system, network, security, and database operations. * Experience with tools such as Tanium or BMC TrueSight Orchestration. * Experience with command-line interfaces (CLI), third-party APIs, and integration. * Server administration experience with Red Hat Enterprise Linux and Windows Server. * Understanding of developing fault-tolerant solutions and knowledge of horizontal scaling and high availability. ## Description * Perform full-stack triaging of alerts to identify the root cause of application performance and stability issues. * Work with product owners and other stakeholders to define and track service level objectives (SLOs). * Design and develop dashboards and reports to communicate key performance metrics. * Identify opportunities to improve alerting posture and update alerts accordingly. * Collaborate with the engineering team to understand application architecture and perform single-point-of-failure analysis. * Create and derive NFR/Workload models to ensure performance and resiliency are considered early in the software development lifecycle. * Execute performance and chaos tests, analyzing results with APM tools to identify stability issues. * Document findings, analysis, and results, presenting them to stakeholders. * Perform analytics on past incidents to understand root causes and implement automation to reduce recurrence. * Demonstrate proficiency with DevOps tools such as JIRA, BMC Remedy, and ServiceNow. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [The Time Paradox: Building Timezone-Safe Python/Django Applications](https://www.wearedevelopers.com/videos/1915-the-time-paradox-building-timezone-safe-python-django-applications) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)