4468 Site Reliability Engineer

Procession Systems
Chantilly, VA, United States
30 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Systems Engineering Bash Shell Cloud Computing Cloud Engineering Software Documentation Information Systems DevOps Distributed Systems Identity and Access Management Python (Programming Language) Linux System Administration
+11 more
Reliability Engineering Data Logging System Availability Infrastructure as Code (IaC) Cloudformation Containerization Infrastructure Automation Frameworks Information Technology Deployment Automation Terraform Docker

Job description

  • Maintain and improve the reliability, availability, and performance of a custom Identity and Access Management (IAM) platform.
  • Monitor production systems and proactively identify, investigate, and resolve infrastructure and application issues.
  • Develop and maintain automation scripts and operational tooling using Python and/or Bash.
  • Support AWS-based infrastructure and services, including deployment, monitoring, and operational management.
  • Collaborate with software engineers, cybersecurity teams, and system administrators to implement reliable, scalable, and secure solutions.
  • Develop monitoring, alerting, logging, and observability capabilities to ensure rapid detection and resolution of operational issues.
  • Participate in incident response, root cause analysis, and post-incident reviews to improve system resiliency.
  • Implement infrastructure improvements using automation and Infrastructure as Code (IaC) best practices where applicable.
  • Create and maintain operational documentation, runbooks, and standard operating procedures.
  • Ensure compliance with DoD security requirements and organizational cybersecurity policies.

Requirements

We are seeking an experienced Site Reliability Engineer (SRE) to support the operations, reliability, and continuous improvement of a mission-critical Identity and Access Management (IAM) platform deployed within a U.S. Department of Defense (DoD) classified environment. This custom-built IAM solution leverages behavioral intelligence and advanced analytics to rapidly identify, detect, and respond to identity- related anomalies and operational issues. The ideal candidate will have strong experience with cloud technologies, automation, infrastructure reliability, and secure system operations. This role requires a proactive engineer who can improve system availability, automate operational tasks, troubleshoot complex issues, and ensure the IAM platform maintains high levels of performance and resilience., * Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related technical discipline.

  • Minimum of 6 years of professional experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, or a related field.
  • Experience supporting applications or infrastructure within Amazon Web Services (AWS).
  • Proficiency in Python and/or Bash scripting for automation and operational tooling.
  • Experience troubleshooting Linux-based systems and distributed applications.
  • Familiarity with monitoring, logging, and observability platforms.
  • Experience supporting mission-critical production environments.
  • Ability to obtain and maintain system documentation and operational procedures., * Experience supporting Identity and Access Management (IAM) platforms.
  • Experience operating systems within classified or highly regulated government environments.
  • Familiarity with Infrastructure as Code tools such as Terraform or AWS CloudFormation.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of container technologies such as Docker and Kubernetes.
  • Understanding of behavioral analytics, identity security, or cybersecurity operations.
  • Strong analytical and problem-solving abilities.
  • Excellent troubleshooting and root cause analysis skills.
  • Effective verbal and written communication skills.
  • Ability to work independently while collaborating across multidisciplinary technical teams.
  • Strong commitment to operational excellence, automation, and continuous improvement.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all