Senior Site Reliability Engineer

Technatomy Corporation
United States
28 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Analysis of Variance (ANOVA) Automation of Tests Bash Shell Cloud Computing Cyber Security DevOps Distributed Systems Elasticsearch Python (Programming Language) Linux System Administration Operational Data Store
+25 more
Performance Tuning Windows PowerShell Role-Based Access Control Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus Zero Trust Network Access Software Engineering Software Vulnerability Management Data Logging Cloud Platform System Grafana Kubernetes Information Technology Deployment Automation Hashicorp Cloudwatch Kibana Terraform Splunk Serverless Computing Docker Golang Programming Languages

Job description

We are seeking an experienced Senior Site Reliability Engineer to serve as a key technical contributor supporting the Technical Director in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role partners with platform, development, operations, monitoring, incident-management, security, and VA stakeholder teams to improve availability, performance, scalability, and operational excellence across mission-critical environments. The Senior Site Reliability Engineer applies software engineering principles to operations while aligning solutions with Federal security and governance requirements., · Partner with the Technical Director to implement and mature Site Reliability Engineering practices across platform services and hosted applications.

· Improve the full service lifecycle from design and deployment through operation and continuous refinement, with a focus on availability, latency, performance, efficiency, and capacity.

· Define, track, and report service-level indicators, service-level objectives, error budgets, and service health measures that guide engineering decisions.

· Build, enhance, and maintain CI/CD pipelines that enable secure, automated, repeatable application and infrastructure delivery.

· Develop and support Infrastructure as Code and configuration automation using Terraform, Ansible, and comparable technologies.

· Integrate automated testing, validation, security checks, rollback, and operational readiness controls into delivery workflows.

· Design and improve monitoring, logging, tracing, alerting, and dashboards to strengthen observability and accelerate issue detection and response.

· Analyze system behavior, performance trends, capacity, failure patterns, and operational data to improve reliability, scalability, and efficiency.

· Reduce operational toil by automating repetitive tasks, improving runbooks, and engineering durable solutions for recurring issues.

· Support AWS infrastructure and Kubernetes, EKS, ECS, Docker, or comparable container platforms with an emphasis on resilience, scalability, and security.

· Contribute to platform modernization, capacity planning, reliability reviews, deployment-pattern improvements, and operational readiness for cloud-native services.

· Implement reliability practices that align with Federal security requirements, including secure configuration, least privilege, vulnerability remediation, and policy-based controls.

· Collaborate with development, platform, operations, monitoring, incident-management, architecture, and cybersecurity teams to improve service and deployment outcomes.

· Participate in incident response, service restoration, root cause analysis, and blameless post-incident reviews for critical systems and services.

· Identify recurring issues, reliability gaps, and failure patterns and drive corrective actions through automation, architecture improvements, and process refinement.

· Strengthen on-call readiness, operational documentation, escalation procedures, and continuous improvement practices that reduce mean time to recovery.

Requirements

· 6+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or related roles supporting enterprise or mission-critical environments.

· Hands-on experience supporting AWS or comparable cloud platforms, Linux-based environments, distributed systems, and production services at scale.

· Strong experience with Infrastructure as Code and configuration automation using Terraform, Ansible, or comparable technologies.

· Experience with Kubernetes, EKS, ECS, Docker, or another container and orchestration platform in production environments.

· Experience building or maintaining CI/CD pipelines and deployment automation for secure, reliable software and infrastructure delivery.

· Strong understanding of monitoring, logging, tracing, observability, incident response, root cause analysis, capacity planning, and performance optimization.

· Proficiency with one or more scripting or programming languages such as Python, Go, Bash, or PowerShell.

· Demonstrated ability to troubleshoot complex systems, automate operational tasks, reduce toil, and implement durable reliability improvements.

· Experience collaborating across software engineering, platform, operations, cybersecurity, architecture, and customer stakeholder teams.

· Strong written and verbal communication skills, including the ability to document technical standards, incidents, risks, operational decisions, and improvement plans.

KNOWLEDGE AND SKILLS DESIRED:

· Experience supporting the Department of Veterans Affairs, another Federal agency, or a regulated enterprise environment with significant security and compliance requirements.

· Experience defining and operationalizing service-level indicators, service-level objectives, error budgets, and production service health metrics.

· Advanced experience with Prometheus, Grafana, CloudWatch, Elasticsearch, Kibana, Splunk, OpenTelemetry, or comparable observability platforms.

· Experience with FedRAMP, NIST, Zero Trust, or other Federal security frameworks relevant to cloud and platform operations.

· Experience supporting healthcare platforms, high-availability enterprise services, or large-scale modernization initiatives.

· Relevant certification such as AWS Certified DevOps Engineer - Professional, AWS Certified Solutions Architect, Certified Kubernetes Administrator, HashiCorp Terraform Associate, or an SRE/DevOps credential.

EDUCATION:

· Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience.

CLEARANCE:

· Must be able to obtain and maintain a Public Trust clearance.

About the company

At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · WWC 2022

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all