> Markdown version of [/jobs/ext/1264737-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1264737-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Technatomy Corporation - **Location:** United States (Remote available) - **Experience:** Starter - **Contract:** Internship / Graduate position - **Skills:** Amazon Web Services, Application Performance Management, Microsoft Azure, Bash Shell, Cloud Computing, Configuration Management, CompTIA Security+, Computer Programming, Continuous Integration, Linux, DevOps, Elasticsearch, Python (Programming Language), Networking Basics, Windows PowerShell, Reliability Engineering, Prometheus, Software Engineering, Data Logging, Scripting, Google Cloud, Enterprise Software Applications, Cloud Platform System, Grafana, Git Flow, Kubernetes, Information Technology, Hashicorp, Cloudwatch, Kibana, Comptia Linux+, Terraform, Splunk, Software Version Control, Docker - **Published:** July 14, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9028521/site-reliability-engineer ## About the Role We are seeking a motivated and detail-oriented Site Reliability Engineer to support the Technical Director's team in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role works with senior engineers, platform and operations teams, and VA stakeholders to support the availability, performance, and operational stability of mission-critical environments. The Site Reliability Engineer applies foundational software engineering and operational practices to improve monitoring, automation, incident response, and service reliability., · 1-3 years of experience in Site Reliability Engineering, DevOps, systems administration, cloud operations, platform support, software engineering, or a related technical role. · Foundational understanding of Linux systems, cloud infrastructure concepts, enterprise application support, and basic networking. · Exposure to scripting or programming using Python, Bash, PowerShell, or a similar language. · Familiarity with monitoring, logging, alerting, troubleshooting, incident response, and service restoration concepts. · Basic knowledge of CI/CD, version control, automation, configuration management, or Infrastructure as Code concepts. · Ability to follow technical procedures, document work accurately, analyze operational information, and escalate issues appropriately. · Strong attention to detail and the ability to learn new cloud, platform, observability, and automation tools quickly. · Ability to work effectively in a collaborative, remote team environment with engineers, operations personnel, and customer stakeholders. KNOWLEDGE AND SKILLS DESIRED: · Internship, academic, lab, or hands-on experience with AWS, Microsoft Azure, Google Cloud, or another cloud platform. · Familiarity with Docker, Kubernetes, EKS, ECS, or another container and orchestration technology. · Exposure to CloudWatch, Grafana, Prometheus, Elasticsearch, Kibana, Splunk, OpenTelemetry, or similar observability tools. · Experience with Git-based workflows, pipeline tooling, or automation through coursework, labs, internships, or professional experience. · Understanding of Federal security, compliance, healthcare technology, or other regulated enterprise environments. · Relevant foundational certification such as AWS Certified Cloud Practitioner, AWS Certified Developer - Associate, CompTIA Linux+, Security+, or HashiCorp Terraform Associate. EDUCATION: · Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience. CLEARANCE: · Must be able to obtain and maintain a Public Trust clearance. ## Description · Support day-to-day Site Reliability Engineering activities across platform services, hosted applications, and cloud environments. · Help maintain service reliability, availability, and performance by following established operational procedures, runbooks, and engineering standards. · Gather and review operational metrics, alerts, logs, and system health information to identify issues and support service improvements. · Maintain monitoring, logging, alerting, and dashboard configurations that improve visibility into infrastructure and application performance. · Participate in incident response, service restoration, escalation, and post-incident follow-up under the guidance of senior team members. · Document incidents, recurring issues, operational procedures, configuration details, and troubleshooting guidance. · Develop simple scripts and automation that reduce manual effort, improve consistency, and address recurring operational tasks. · Support CI/CD processes and environment maintenance for application and infrastructure delivery across development, test, and production environments. · Assist with Infrastructure as Code, configuration changes, and environment updates using approved tools, templates, and team guidance. · Perform routine operational checks and support activities for AWS and container-based platforms. · Maintain service inventory, configuration records, operational documentation, and other artifacts used by the reliability team. · Assist with validation, testing, deployment readiness, and operational acceptance activities for releases and environment changes. · Follow established security, access, change, and operational procedures that support Federal compliance and secure administration. · Collaborate with software, infrastructure, platform, monitoring, incident-management, and support teams to resolve issues and improve reliable service delivery. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)