> Markdown version of [/jobs/ext/3038500-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3038500-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Vital Tech Solutions - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, HP Systems Insight Manager, Python (Programming Language), Linux Commands, Reliability Engineering, Ansible, Data Streaming, Cloud Platform System, Delivery Pipeline, Apache Spark, Reliability of Systems, Git, Information Technology, Terraform, Software Version Control, Docker, Databricks - **Published:** September 23, 2026 - **Apply:** https://www.thejobnetwork.com/job/f7efc359-f645-4184-8022-188738ca7b76/senior-site-reliability-engineer-sre ## About the Role * Active Secret security clearance or higher is required * Strong experience supporting and maintaining production infrastructure. * Hands-on Python experience within operational, infrastructure, or support environments. * Professional experience with Terraform, Ansible, and Docker. * Experience supporting CI/CD pipelines and deployment workflows. * Strong proficiency with Git and version-control practices. * Strong Linux command-line and systems operations experience. * Experience monitoring, diagnosing, and troubleshooting production systems. * Strong incident-response and root-cause analysis capabilities. * Ability to anticipate and resolve complex operational issues. * Strong communication skills and the ability to work effectively in a collaborative, client-facing environment. * Ability to learn and adapt to new technologies quickly. * Ability to work East Coast business hours. Preferred Qualifications * Experience operating infrastructure within AWS or another major cloud platform. * Hands-on experience with AWS services including EKS, S3, and EMR. * Familiarity with Spark, JupyterHub, and Hue in an operational environment. * Databricks experience. * Experience supporting federal government, regulated, or other security-sensitive environments. * Experience working directly with external clients or government stakeholders. * Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline. Full Time ## Description This role is focused on the reliable, secure, and efficient operation of critical production infrastructure. The Senior Site Reliability Engineer will be responsible for maintaining production environments, troubleshooting incidents, supporting deployments, improving monitoring and diagnostics, and driving greater uptime and operational stability., * Maintain, monitor, and troubleshoot production environments to support system uptime, reliability, and performance. * Manage and operate infrastructure using Terraform, Ansible, and Docker. * Support and maintain CI/CD pipelines and automate operational workflows using Git and related tooling. * Ensure the ongoing reliability of systems operating within AWS environments, including EKS, S3, and EMR. * Support operational use of technologies such as Spark, JupyterHub, and Hue. * Diagnose and resolve infrastructure and application issues. * Conduct root-cause analysis and drive long-term resolution of recurring incidents. * Implement, refine, and maintain infrastructure and application monitoring, alerting, and diagnostics. * Support deployment activities, maintenance windows, and production changes. * Optimize data flows and storage integrations. * Collaborate with engineering, product, and client stakeholders to communicate issues, coordinate maintenance, and support operational priorities. * Contribute to continuous improvement of operational processes, platform documentation, and reliability best practices. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)