> Markdown version of [/jobs/ext/2196556-senior-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2196556-senior-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Reliability Engineer - **Company:** Barbaricum Llc - **Location:** Washington, DC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Amazon Web Services, Systems Engineering, Microsoft Azure, Bash Shell, Cloud Computing, Software Documentation, Program Optimization, Cyber Security, System Configuration, Monitoring of Systems, Python (Programming Language), Network Troubleshooting, Linux System Administration, Automation of Marketing, Performance Tuning, Windows PowerShell, Reliability Engineering, Site Reliability Engineering Practices, Cloud Services, Ansible, Runbook, Scripting, Google Cloud, Cloud Platform System, Reliability of Systems, Infrastructure Automation Frameworks, Information Technology, Performance Monitor, Puppet, Devsecops - **Published:** August 23, 2026 - **Apply:** https://www.clearancejobs.com/jobs/8983258/senior-reliability-engineer ## About the Role * Expert knowledge of site reliability engineering practices, system monitoring, incident management, automation, performance tuning, and operational resilience. * Strong understanding of Windows and Linux administration, infrastructure operations, system configuration, service management, and troubleshooting practices. * Experience with automation platforms and configuration management tools such as Ansible, Puppet, Chef, or similar technologies. * Proficiency with scripting languages such as Python, Shell, PowerShell, or similar tools used to automate operational and infrastructure tasks. * Knowledge of cloud services and infrastructure across AWS, Microsoft Azure, Google Cloud, or comparable cloud environments. * Strong understanding of network troubleshooting, configuration, connectivity analysis, system dependencies, and performance bottleneck identification. * Ability to design, interpret, and maintain dashboards, alerts, metrics, logs, and operational reporting that support service health and decision-making. * Ability to conduct root cause analysis, post-incident reviews, and corrective action planning in complex technical environments. * Strong problem-solving skills and the ability to work under pressure during outages, impairments, and time-sensitive operational issues. * Excellent written and verbal communication skills, with the ability to explain technical findings, incident impacts, and reliability recommendations to technical and non-technical stakeholders., * Bachelor's degree in Computer Science, Information Technology, Systems Engineering, Cybersecurity, or a related field; Master's degree preferred. * Certifications related to cloud computing, system administration, site reliability engineering, DevSecOps, or automation are beneficial. * 10+ years of experience in site reliability engineering, systems administration, infrastructure operations, cloud operations, DevSecOps, or a similar technical role, particularly in a government, federal, defense, or secure IT setting. * Demonstrated experience maintaining reliable, scalable, and efficiently managed IT systems across on-premises, cloud, or hybrid environments. * Experience developing automated infrastructure, operational scripts, monitoring solutions, dashboards, runbooks, and configuration standards. * Experience supporting incident response, system outage resolution, post-incident reviews, root cause analysis, and operational improvement initiatives. * Experience collaborating with development, infrastructure, cloud, cybersecurity, and program teams to improve reliability, security, and service performance. * DoD Secret Security Clearance. ## Description * Monitor and maintain system reliability, availability, and performance across on-premises, cloud, and hybrid IT environments supporting MC&FP mission requirements. * Implement proactive performance monitoring, automated alerting, incident response workflows, and resilience engineering practices to reduce downtime and improve operational visibility. * Develop, maintain, and improve scalable automated infrastructure solutions that support reliable system operations and repeatable service delivery. * Implement rollback strategies, recovery approaches, and chaos engineering practices to validate resilience, reduce operational risk, and improve system stability. * Analyze usage patterns, capacity trends, and performance indicators to support dynamic scaling, resource optimization, and system improvement decisions. * Develop and maintain real-time operational dashboards, reports, and metrics that enable rapid decision-making, leadership awareness, and system optimization. * Respond to and resolve system outages, impairments, and service disruptions while coordinating with technical teams to minimize mission impact. * Conduct post-incident reviews to identify root causes, document lessons learned, and implement preventative measures that reduce recurrence. * Collaborate with software developers, cloud engineers, cybersecurity personnel, and operations teams to improve services, reliability patterns, deployment practices, and operational standards. * Create and maintain system documentation, configuration standards, operational runbooks, monitoring procedures, and service reliability guidance. * Automate common operations tasks to reduce manual workloads, improve consistency, and increase system efficiency. * Implement security best practices across operational activities, infrastructure automation, monitoring, incident response, and system administration functions. ## Related Videos - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)