> Markdown version of [/jobs/ext/1188285-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1188285-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** AdaptiveVets Solutions, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $113,050.0 - $133,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Amazon Web Services, Cloud Computing, CompTIA Security+, Monitoring of Systems, Reliability Engineering, Data Logging, Mttr, Kubernetes, Information Technology, Splunk, Dynatrace - **Published:** July 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ea188816ea79b13a ## About the Role * Bachelor's degree in Computer Science, Engineering, IT, or related field (4 additional years of relevant experience may substitute) * 5+ years of site reliability engineering or cloud operations experience in Kubernetes-based environments * Hands-on experience with Dynatrace, Splunk, and enterprise monitoring and alerting platforms * Experience with AWS GovCloud operations, Kubernetes/EKS node management, and automated recovery procedures * Ability to support 24x7 on-call operations and respond rapidly to Critical incidents * Proficiency with SRE automation scripting and root cause analysis methodologies Preferred Qualifications * Experience with SRE frameworks including SLI/SLO/SLA definition and error budget management * AWS Certified SysOps Administrator or equivalent * Certified Kubernetes Administrator (CKA) Education: * Bachelor's degree in Computer Science, Engineering, or IT (Required) Experience: * Site reliability engineering in cloud Kubernetes environments: 5 years (Required) License/Certification: * AWS Certified SysOps Administrator (Preferred) * Certified Kubernetes Administrator (CKA) (Preferred) * CompTIA Security+ (Preferred) Location and Ability to Commute: * Remote Security clearance: * Public Trust (Preferred), * Bachelor's (Required) Experience: * Agile: 4 years (Preferred) * Site reliability engineering in cloud Kubernetes environment: 5 years (Required) Language: * English (Required) License/Certification: * AWS Certified SysOps Administrator (Preferred) * Certified Kubernetes Administrator (CKA) (Preferred) * CompTIA Security+ (Preferred) ## Description * Implement and manage platform observability using Dynatrace and Splunk, including monitoring, alerting, centralized logging, and operational dashboards. * Support incident response operations, coordinating with the Monitoring and Incident Management Manager to detect, escalate, and resolve incidents within required SLA timeframes. * Implement and maintain automated recovery procedures including node failure detection, cordon/drain/replace workflows, and failover automation. * Participate in 24x7 on-call rotation to maintain platform availability and respond rapidly to Critical and High severity incidents. * Proactively engage in PI planning alongside supported application teams to address infrastructure dependencies and capacity planning. * Develop and maintain operational runbooks, SRE automation scripts, and post-incident root cause analysis documentation. * Monitor resource saturation metrics and autoscaling policy effectiveness, alerting before defined thresholds are exceeded. * Track and report SRE metrics including availability, incident frequency, MTTR, and reliability trends. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [It's Not Vibe Coding If You Know What You're Doing](https://www.wearedevelopers.com/videos/100119-it-s-not-vibe-coding-if-you-know-what-you-re-doing) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)