> Markdown version of [/jobs/ext/2116950-cloud-sre-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2116950-cloud-sre-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud SRE Systems Engineer - **Company:** Integral Consulting Services - **Location:** Tysons, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Systems Engineering, JIRA, Microsoft Azure, Cloud Computing, Fault Tolerance, Github, Reliability Engineering, Site Reliability Engineering Practices, Software Deployment, Backup and Restore, Data Logging, HybridCloud, Information Technology, Performance Monitor, Microservices - **Published:** August 19, 2026 - **Apply:** https://www.jofdav.com/jobs/59313220-cloud-sre-systems-engineer ## About the Role * Bachelor's Degree in computer science or IT related degree with 7-10 years' experience * Knowledge of modern SRE practices: automation, resilience engineering, observability, fault tolerance, SLT/SLA adherence, and incident analysis. * Experience producing monitoring reports, SLT assessments, postmortems, and operational dashboards * Experience with Jira and GitHub * Competency maintaining IRPs, DRPs, backup/restore plans, and executing continuity operations * Strong documentation and communication skills for reporting outages, SLT performance, and incident updates * Public Trust Preferred: * Experience with AWS VA Enterprise Cloud (VAEC) * Experience with Azure DevOps ## Description The Cloud Site Reliability Engineering (SRE) Systems Engineer is responsible for ensuring reliability, availability, performance, and operational excellence of VA product environments deployed in the VA Enterprise Cloud (VAEC) for the Department of Veterans Affairs (VA), Office of Information and Technology (OIT), Product Delivery Services (PDS), Benefits and Memorials (BAM), Memorial and Benefit Services (MBS) deliver secure, reliable, and effective Information Technology (IT) solutions that support the Department's mission., * Applies SRE practices, manages cloud environments, performs patching and automation, ensures compliance with required Service Level Targets (SLTs), supports deployments, handles incidents, and maintains monitoring, observability, and operational documentation * Maintain all environments in an operational state across VAEC, including application/system hardware 24x7x365. * Apply OS and application patches and perform backup and restoration activities. * Support software releases (often outside core hours) and validate maximum load after production deployment. * Manage and optimize cloud resource utilization, including capacity planning and forecasting * Implement SRE practices and patterns to improve reliability, performance, automation, fault tolerance, telemetry, logging, and alerting across distributed and microservices architectures. * Improve fault tolerance for distributed/microservices systems when components or resources become temporarily unavailable. * Conduct trend analysis of system performance and resource utilization to forecast future demand and identify bottlenecks. * Implement realtime proactive monitoring dashboards enabling visibility into system health and alerting. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)