> Markdown version of [/jobs/ext/1459842-4468-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1459842-4468-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # 4468 Site Reliability Engineer - **Company:** Procession Systems - **Location:** Chantilly, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Systems Engineering, Bash Shell, Cloud Computing, Cloud Engineering, Software Documentation, Information Systems, DevOps, Distributed Systems, Identity and Access Management, Python (Programming Language), Linux System Administration, Reliability Engineering, Data Logging, System Availability, Infrastructure as Code (IaC), Cloudformation, Containerization, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Terraform, Docker - **Published:** July 27, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9058607/4468-site-reliability-engineer ## About the Role We are seeking an experienced Site Reliability Engineer (SRE) to support the operations, reliability, and continuous improvement of a mission-critical Identity and Access Management (IAM) platform deployed within a U.S. Department of Defense (DoD) classified environment. This custom-built IAM solution leverages behavioral intelligence and advanced analytics to rapidly identify, detect, and respond to identity- related anomalies and operational issues. The ideal candidate will have strong experience with cloud technologies, automation, infrastructure reliability, and secure system operations. This role requires a proactive engineer who can improve system availability, automate operational tasks, troubleshoot complex issues, and ensure the IAM platform maintains high levels of performance and resilience., * Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical discipline. * Minimum of 6 years of professional experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, or a related field. * Experience supporting applications or infrastructure within Amazon Web Services (AWS). * Proficiency in Python and/or Bash scripting for automation and operational tooling. * Experience troubleshooting Linux-based systems and distributed applications. * Familiarity with monitoring, logging, and observability platforms. * Experience supporting mission-critical production environments. * Ability to obtain and maintain system documentation and operational procedures., * Experience supporting Identity and Access Management (IAM) platforms. * Experience operating systems within classified or highly regulated government environments. * Familiarity with Infrastructure as Code tools such as Terraform or AWS CloudFormation. * Experience with CI/CD pipelines and deployment automation. * Knowledge of container technologies such as Docker and Kubernetes. * Understanding of behavioral analytics, identity security, or cybersecurity operations. * Strong analytical and problem-solving abilities. * Excellent troubleshooting and root cause analysis skills. * Effective verbal and written communication skills. * Ability to work independently while collaborating across multidisciplinary technical teams. * Strong commitment to operational excellence, automation, and continuous improvement. ## Description * Maintain and improve the reliability, availability, and performance of a custom Identity and Access Management (IAM) platform. * Monitor production systems and proactively identify, investigate, and resolve infrastructure and application issues. * Develop and maintain automation scripts and operational tooling using Python and/or Bash. * Support AWS-based infrastructure and services, including deployment, monitoring, and operational management. * Collaborate with software engineers, cybersecurity teams, and system administrators to implement reliable, scalable, and secure solutions. * Develop monitoring, alerting, logging, and observability capabilities to ensure rapid detection and resolution of operational issues. * Participate in incident response, root cause analysis, and post-incident reviews to improve system resiliency. * Implement infrastructure improvements using automation and Infrastructure as Code (IaC) best practices where applicable. * Create and maintain operational documentation, runbooks, and standard operating procedures. * Ensure compliance with DoD security requirements and organizational cybersecurity policies. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)