> Markdown version of [/jobs/ext/1264087-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1264087-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Technatomy Corporation - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Analysis of Variance (ANOVA), Automation of Tests, Bash Shell, Cloud Computing, Cyber Security, DevOps, Distributed Systems, Elasticsearch, Python (Programming Language), Linux System Administration, Operational Data Store, Performance Tuning, Windows PowerShell, Role-Based Access Control, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Zero Trust Network Access, Software Engineering, Software Vulnerability Management, Data Logging, Cloud Platform System, Grafana, Kubernetes, Information Technology, Deployment Automation, Hashicorp, Cloudwatch, Kibana, Terraform, Splunk, Serverless Computing, Docker, Golang, Programming Languages - **Published:** July 14, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9028520/senior-site-reliability-engineer ## About the Role · 6+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or related roles supporting enterprise or mission-critical environments. · Hands-on experience supporting AWS or comparable cloud platforms, Linux-based environments, distributed systems, and production services at scale. · Strong experience with Infrastructure as Code and configuration automation using Terraform, Ansible, or comparable technologies. · Experience with Kubernetes, EKS, ECS, Docker, or another container and orchestration platform in production environments. · Experience building or maintaining CI/CD pipelines and deployment automation for secure, reliable software and infrastructure delivery. · Strong understanding of monitoring, logging, tracing, observability, incident response, root cause analysis, capacity planning, and performance optimization. · Proficiency with one or more scripting or programming languages such as Python, Go, Bash, or PowerShell. · Demonstrated ability to troubleshoot complex systems, automate operational tasks, reduce toil, and implement durable reliability improvements. · Experience collaborating across software engineering, platform, operations, cybersecurity, architecture, and customer stakeholder teams. · Strong written and verbal communication skills, including the ability to document technical standards, incidents, risks, operational decisions, and improvement plans. KNOWLEDGE AND SKILLS DESIRED: · Experience supporting the Department of Veterans Affairs, another Federal agency, or a regulated enterprise environment with significant security and compliance requirements. · Experience defining and operationalizing service-level indicators, service-level objectives, error budgets, and production service health metrics. · Advanced experience with Prometheus, Grafana, CloudWatch, Elasticsearch, Kibana, Splunk, OpenTelemetry, or comparable observability platforms. · Experience with FedRAMP, NIST, Zero Trust, or other Federal security frameworks relevant to cloud and platform operations. · Experience supporting healthcare platforms, high-availability enterprise services, or large-scale modernization initiatives. · Relevant certification such as AWS Certified DevOps Engineer - Professional, AWS Certified Solutions Architect, Certified Kubernetes Administrator, HashiCorp Terraform Associate, or an SRE/DevOps credential. EDUCATION: · Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience. CLEARANCE: · Must be able to obtain and maintain a Public Trust clearance. ## Description We are seeking an experienced Senior Site Reliability Engineer to serve as a key technical contributor supporting the Technical Director in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role partners with platform, development, operations, monitoring, incident-management, security, and VA stakeholder teams to improve availability, performance, scalability, and operational excellence across mission-critical environments. The Senior Site Reliability Engineer applies software engineering principles to operations while aligning solutions with Federal security and governance requirements., · Partner with the Technical Director to implement and mature Site Reliability Engineering practices across platform services and hosted applications. · Improve the full service lifecycle from design and deployment through operation and continuous refinement, with a focus on availability, latency, performance, efficiency, and capacity. · Define, track, and report service-level indicators, service-level objectives, error budgets, and service health measures that guide engineering decisions. · Build, enhance, and maintain CI/CD pipelines that enable secure, automated, repeatable application and infrastructure delivery. · Develop and support Infrastructure as Code and configuration automation using Terraform, Ansible, and comparable technologies. · Integrate automated testing, validation, security checks, rollback, and operational readiness controls into delivery workflows. · Design and improve monitoring, logging, tracing, alerting, and dashboards to strengthen observability and accelerate issue detection and response. · Analyze system behavior, performance trends, capacity, failure patterns, and operational data to improve reliability, scalability, and efficiency. · Reduce operational toil by automating repetitive tasks, improving runbooks, and engineering durable solutions for recurring issues. · Support AWS infrastructure and Kubernetes, EKS, ECS, Docker, or comparable container platforms with an emphasis on resilience, scalability, and security. · Contribute to platform modernization, capacity planning, reliability reviews, deployment-pattern improvements, and operational readiness for cloud-native services. · Implement reliability practices that align with Federal security requirements, including secure configuration, least privilege, vulnerability remediation, and policy-based controls. · Collaborate with development, platform, operations, monitoring, incident-management, architecture, and cybersecurity teams to improve service and deployment outcomes. · Participate in incident response, service restoration, root cause analysis, and blameless post-incident reviews for critical systems and services. · Identify recurring issues, reliability gaps, and failure patterns and drive corrective actions through automation, architecture improvements, and process refinement. · Strengthen on-call readiness, operational documentation, escalation procedures, and continuous improvement practices that reduce mean time to recovery. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)