> Markdown version of [/jobs/ext/2597832-data-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2597832-data-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Site Reliability Engineer (SRE) - **Company:** General Dynamics Information Technology - **Location:** Fairfax, VA, United States (Remote available) - **Experience:** Expert - **Salary:** $111,155.0 - $150,385.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Program Optimization, Computer Programming, Computer Networks, Databases, Continuous Integration, Data as a Services, Data Governance, Data Infrastructure, Data Transformation, DevOps, Disaster Recovery, Multi-Factor Authentication, Fault Tolerance, Github, Identity and Access Management, Integrated Development Environments, Python (Programming Language), Machine Learning, Automation of Marketing, Octopus Deploy, Performance Tuning, Windows PowerShell, Systems Development Life Cycle, Reliability Engineering, Cloud Services, Prometheus, Security Information and Event Management, Single Sign-On, Software Engineering, Toolchain, Virtual Machines, Beta Testing, Datadog, Data Logging, Scripting, Enterprise Software Applications, Cloud Platform System, Cloud Monitoring, DevOps Tools - Open-source, Grafana, Reliability of Systems, Gitlab, Cloudformation, Build Management, Containerization, Infrastructure Automation Frameworks, Information Technology, Bicep, Data Management, Cloudwatch, Elastic Beanstalk, Terraform, New Relic (SaaS), Software Version Control, Devsecops, Jenkins, Static Application Security Testing, Dynamic Application Security Testing - **Published:** August 27, 2026 - **Apply:** https://dejobs.org/x/x/87AF165F96E54C779B56021C6322DB86/job/ ## About the Role CI/CD,Containerization,Structured Query Language (SQL) Development Experience: 5 + years of related experience, * Education: Bachelor's degree in Computer Science, Software Engineering, or related field. (Or equivalent experience.) * Experience: 5+ years' experience in IT systems engineering, systems development, systems coding, and programming. * Deep expertise with AWS services, including monitoring, logging, compute, storage, and networking. * Proficiency in Infrastructure as Code (IaC) tools like Terraform, AWS CloudFormation, or Azure Bicep. * Hands-on experience with monitoring and APM tools such as CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, New Relic, etc. * Solid understanding of incident response, change management, and ITIL-based operational support. * Familiarity with CI/CD toolchains and automation platforms (Jenkins, GitHub Actions, GitLab, ArgoCD). * Strong scripting skills (Python, PowerShell, Bash) for automation and orchestration. * Advanced experience in providing DevSecOps implementation using GitOps, or similar tools. * Experienced in developing, testing, and maintaining containerized applications. * Expert knowledge of source version control, build/release tools and methodologies, CI/CD pipelines and the Software Build process. * Experience in building and maintaining CI/CD pipelines for large enterprises that consist of a large number of complex applications. * Ability to be flexible and work on several different products while supporting multiple teams * Experience with FinOps practices, cost modeling, forecasting, and optimization tools within cloud platforms. * Understanding of federal compliance and security frameworks (e.g., FedRAMP, NIST, JISF Rev 5). * Ability to analyze logs and metrics and conduct performance tuning for cloud-based services and applications. * Experience working across multiple product teams to get a grasp of a product and/or programs overall state of health. * ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are a plus. COMMUNICATION & ORGANIZATIONAL SKILLS * Excellent presentation and communication skills. * Consultant mindset with the ability to work with high level customer stakeholders and build excellent customer relationships. * Experience identifying and applying industry tools, solutions, methods best practices, and emerging technologies. * Strong analytical skills and problem-solving skills with the ability to formulate and communicate recommendations for improvement. * Experience with process design and documentation methodologies, and design and production of quality deliverables, process and use case modeling, business case development. * Demonstrated ability to work effectively, independently, and as part of a team. Security Clearance Level: Must be able to pass a background check to obtain a position of Public Trust. Must be a US Person (Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen). ## Description Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States. GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Data Site Reliability Engineer (SRE) will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program. The successful candidate will be responsible for providing technical leadership for the day-to-day operational support, reliability, performance, and continuous improvement of the CMM data platforms, pipelines, applications, and analytics services. This role ensures that data services remain secure, available, reliable, and aligned with established service levels, data governance standards, architecture principles, and operational procedures., * Provide comprehensive real-time monitoring, incident and event management, capacity planning, and operational reporting to support application deployments, maintain system health, predict demand, and align cloud operations with evolving business and security objectives. * Maintain and audit user roles and responsibilities in cloud environments. * Integrate Single Sign On (SSO), Multi-Factor Authentication (MFA) and group identity management managed through the Judiciary Enterprise Network Information Exchange (JENIE) for enforcing least privilege access. * Adhere to guidelines prescribed by the Government and continuously assess and improve credential management processes for all user credentials. * Provide Disaster Recovery (DR) and Continuity of Operations (COOP) options. This must include high-availability options, including fault-tolerant and automated failover designs. * Integrate DevSecOps tools and processes seamlessly with enterprise systems (Integrated Development Environments (IDEs), ticketing, monitoring, etc.) to avoid fragmentation and ensure unified security posture. * Provide and manage a centralized secrets management system with automated rotation, access logging, and policy enforcement to securely store, manage, and control access to sensitive information and to prevent unauthorized access and data breaches for any administrative user account. * Integrate security tools (example: SAST, DAST, SCA, CSPM) into pipelines for continuous assessment and remediation. * Implement unified, automated, continuous monitoring (24/7/365) systems and tools for security, performance, and compliance across all environments, leveraging dashboards and alerting for real-time visibility. Provide supplemental monitoring of event response activities beyond normal business hours (7a.m - 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts. * Ensure automated generation and management of Software Bill of Materials (SBOM) for all deployed artifacts, supporting transparency and compliance. * Provide diagnostics, metrics' gathering, and performance tuning services. * Provide canary release function for end-user testing to support beta testing. * Configure an alert mechanism so that the support teams can react in an instance of unusual behavior. * Implement and operate a comprehensive incident and event management process, including integration with enterprise SIEM solutions, automated alerting, escalation workflows, and root cause analysis for all critical incidents. * Provide engineering support to ensure prompt detection, logging, diagnosis, escalation, and resolution of incidents to restore normal service operations as quickly as possible and minimize impact. * Perform systems support in identifying, analyzing, and eliminating the root causes of recurring incidents to minimize continued adverse impacts and potential degradation of services. * Make recommendations for the improvement of Incident and Problem management consistent with industry's best practices for the cloud. * Maintain knowledge base of known issues, resolutions, and best practices for operational continuity. * Perform automated health checks across the full stack (Operating System, Application, Database and PaaS services) at agreed levels on an agreed frequency. * Provide a monthly issues management report. The report shall include cloud-related incidents, any stability and performance issues, configurations issues, quantity of tickets received, and time duration to resolve tickets. * Develop and implement thresholds, rules, and response procedures based on product team's recommendation. * Monitor resource utilization (e.g., CPU, Memory, Disk Space) for the cloud hosted Virtual Machines (VMs) and other cloud services. * Manage the resolution procedures for any threshold breaches for cloud resources. * Improves system reliability, observability, automation, scalability, and operational resilience through engineering practices. * Monitors, maintains, and optimizes cloud infrastructure, databases, and platform services for reliability and performance. * Act as FinOps Analyst and perform cost optimization. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)