> Markdown version of [/jobs/ext/2695933-cloud-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2695933-cloud-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Site Reliability Engineer - **Company:** General Dynamics Information Technology - **Location:** Fairfax, VA, United States - **Experience:** Expert - **Salary:** $128,039.0 - $173,229.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Software Applications, Microsoft Azure, Bash Shell, Cloud Computing, Program Optimization, Continuous Integration, Data Governance, Data Infrastructure, Data Transformation, DevOps, Github, Python (Programming Language), Machine Learning, Automation of Marketing, Octopus Deploy, Platform as a Service (PAAS), Performance Tuning, Windows PowerShell, Systems Development Life Cycle, Reliability Engineering, Cloud Services, Ansible, Prometheus, Runbook, Security Information and Event Management, Software Engineering, Toolchain, Datadog, Data Logging, Scripting, Cloud Monitoring, Grafana, Infrastructure as Code (IaC), Gitlab, Cloudformation, Build Management, Containerization, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Bicep, Machine Learning Operations, Cloudwatch, Terraform, New Relic (SaaS), Software Version Control, Devsecops, Jenkins - **Published:** September 3, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3375521740&tx=DT10394UYZ&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role Ansible (Software),Containerization,Information Technology (IT) Experience: 5 + years of related experience, * Education: Technical Training, Certification(s) or Degree required; Bachelor's Degree (or equivalent experience) in Computer Science, Software Engineering, or related field strongly preferred * Experience: 5+ years experience in IT system engineering, systems development, systems coding and programming required * Deep expertise with AWS, Azure, or GCP services, including monitoring, logging, compute, storage, and networking * Proficiency in Infrastructure as Code (IaC) tools like Terraform, AWS CloudFormation, or Azure Bicep * Hands-on experience with monitoring and APM tools such as CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, New Relic, etc. * Solid understanding of incident response, change management , and ITIL-based operational support * Familiarity with CI/CD toolchains and automation platforms (Jenkins, GitHub Actions, GitLab, ArgoCD) * Strong scripting skills (Python, PowerShell, Bash) for automation and orchestration * Advanced experience in providing DevSecOps implementation using GitOps, or similar tools * Experienced in developing, testing, and maintaining containerized applications * Expert knowledge of source version control, build/release tools and methodologies, CI/CD pipelines and the Software Build process * Experience in building and maintaining CI/CD pipelines for large enterprises that consists of a large number of complex applications * Ability to be flexible and work on several different products while supporting multiple teams * Experience with FinOps practices , cost modeling, forecasting, and optimization tools within cloud platforms * Understanding of federal compliance and security frameworks (e.g., FedRAMP, NIST, JISF Rev 5) * Ability to analyze logs and metrics and conduct performance tuning for cloud-based services and applications * Experience working across multiple product teams to get a grasp of a product and/or programs overall state of health * ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are a plus Communication and Organizational: * Excellent presentation and communication skills * Consultant mindset with the ability to work with high level customer stakeholders and build excellent customer relationship * Experience identifying and applying industry tools, solutions, methods best practices, and emerging technologies * Strong analytical skills and problem-solving skills with the ability to formulate and communicate recommendations for improvement * Experience with process design and documentation methodologies, and design and production of quality deliverables, process and use case modeling, business case development * Demonstrated ability to work effectively, independently, and as part of a team Security Clearance Level: Must be able to pass a background check to obtain a position of Public Trust. Must be a US Person (Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen). ## Description Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States. GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Cloud Site Reliability Engineer will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program. They will work as part of an agile development team to help build and support the modernization of enterprise-class software applications., * Ensure operational stability, availability, performance, and scalability of cloudhosted systems across production and development environments supporting multiple agile teams * Provide real-time monitoring , alerting, incident response, and health checks for infrastructure and applications across all cloud layers (OS, app, DB) * Implement and maintain dashboards, visualizations, and reports for system health, event management, and cost optimization using native CSP tools * Manage cloud resource thresholds and automate capacity planning , forecasting, and resource optimization strategies * Perform incident and event management (SIEM) operations, and support issue diagnosis, resolution, and reporting including RCA documentation * Track, document, and report monthly issues , including system performance, stability, ticket volumes, and time-to-resolution metrics * Monitor resource utilization (CPU, memory, disk space) across all deployed VMs, containers, and PaaS components * Provide supplemental monitoring of event response activities beyond normal business hours (7a.m - 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts. * Contribute to the implementation of the Enterprise FinOps framework , including forecasting, budget control, and right-sizing analysis * Validate billing accuracy and support TCO analysis across cloud service usage * Support deployment automation and ensure systems are resilient, repeatable, and scalable via Infrastructure as Code (IaC) * Integrate operations with DevSecOps , MLOps, and CI/CD pipelines for seamless deployment and management * Execute daily or agreed frequency system health checks and maintain operational Runbooks and SOPs * Ensures that DevSecOps principals are followed through the entire software delivery lifecycle ## Related Videos - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)