Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+23 more
Job description
· Support day-to-day Site Reliability Engineering activities across platform services, hosted applications, and cloud environments.
· Help maintain service reliability, availability, and performance by following established operational procedures, runbooks, and engineering standards.
· Gather and review operational metrics, alerts, logs, and system health information to identify issues and support service improvements.
· Maintain monitoring, logging, alerting, and dashboard configurations that improve visibility into infrastructure and application performance.
· Participate in incident response, service restoration, escalation, and post-incident follow-up under the guidance of senior team members.
· Document incidents, recurring issues, operational procedures, configuration details, and troubleshooting guidance.
· Develop simple scripts and automation that reduce manual effort, improve consistency, and address recurring operational tasks.
· Support CI/CD processes and environment maintenance for application and infrastructure delivery across development, test, and production environments.
· Assist with Infrastructure as Code, configuration changes, and environment updates using approved tools, templates, and team guidance.
· Perform routine operational checks and support activities for AWS and container-based platforms.
· Maintain service inventory, configuration records, operational documentation, and other artifacts used by the reliability team.
· Assist with validation, testing, deployment readiness, and operational acceptance activities for releases and environment changes.
· Follow established security, access, change, and operational procedures that support Federal compliance and secure administration.
· Collaborate with software, infrastructure, platform, monitoring, incident-management, and support teams to resolve issues and improve reliable service delivery.
Requirements
We are seeking a motivated and detail-oriented Site Reliability Engineer to support the Technical Director’s team in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role works with senior engineers, platform and operations teams, and VA stakeholders to support the availability, performance, and operational stability of mission-critical environments. The Site Reliability Engineer applies foundational software engineering and operational practices to improve monitoring, automation, incident response, and service reliability., · 1-3 years of experience in Site Reliability Engineering, DevOps, systems administration, cloud operations, platform support, software engineering, or a related technical role.
· Foundational understanding of Linux systems, cloud infrastructure concepts, enterprise application support, and basic networking.
· Exposure to scripting or programming using Python, Bash, PowerShell, or a similar language.
· Familiarity with monitoring, logging, alerting, troubleshooting, incident response, and service restoration concepts.
· Basic knowledge of CI/CD, version control, automation, configuration management, or Infrastructure as Code concepts.
· Ability to follow technical procedures, document work accurately, analyze operational information, and escalate issues appropriately.
· Strong attention to detail and the ability to learn new cloud, platform, observability, and automation tools quickly.
· Ability to work effectively in a collaborative, remote team environment with engineers, operations personnel, and customer stakeholders.
KNOWLEDGE AND SKILLS DESIRED:
· Internship, academic, lab, or hands-on experience with AWS, Microsoft Azure, Google Cloud, or another cloud platform.
· Familiarity with Docker, Kubernetes, EKS, ECS, or another container and orchestration technology.
· Exposure to CloudWatch, Grafana, Prometheus, Elasticsearch, Kibana, Splunk, OpenTelemetry, or similar observability tools.
· Experience with Git-based workflows, pipeline tooling, or automation through coursework, labs, internships, or professional experience.
· Understanding of Federal security, compliance, healthcare technology, or other regulated enterprise environments.
· Relevant foundational certification such as AWS Certified Cloud Practitioner, AWS Certified Developer - Associate, CompTIA Linux+, Security+, or HashiCorp Terraform Associate.
EDUCATION:
· Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience.
CLEARANCE:
· Must be able to obtain and maintain a Public Trust clearance.
About the company
At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.clearancejobs.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Highest Paying Tech Companies for Developers
Find a Developer Job: 12 Best Job Sites For Developers
What Are The Top Skills Required For Azure Developers?