Cloud SRE Systems Engineer

Integral Federal
United States
11 days ago
Apply on jobs.military.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

JIRA Fault Tolerance Github Site Reliability Engineering Practices Information Technology

Requirements

  • Bachelor’s Degree in computer science or IT related degree with 7-10 years’ experience\n
  • Knowledge of modern SRE practices: automation, resilience engineering, observability, fault tolerance, SLT/SLA adherence, and incident analysis. \n
  • Experience producing monitoring reports, SLT assessments, postmortems, and operational dashboards\n
  • Experience with Jira and GitHub\n
  • Competency maintaining IRPs, DRPs, backup/restore plans, and executing continuity operations\n
  • Strong documentation and communication skills for reporting outages, SLT performance, and incident updates\n
  • Public Trust\n

Benefits & conditions

The Cloud Site Reliability Engineering (SRE) Systems Engineer is responsible for ensuring reliability, availability, performance, and operational excellence of VA product environments deployed in the VA Enterprise Cloud (VAEC) for the Department of Veterans Affairs (VA), Office of Information and Technology (OIT), Product Delivery Services (PDS), Benefits and Memorials (BAM), Memorial and Benefit Services (MBS) deliver secure, reliable, and effective Information Technology (IT) solutions that support the Department’s mission.\n \n \nResponsibilities:\n \n

  • Applies SRE practices, manages cloud environments, performs patching and automation, ensures compliance with required Service Level Targets (SLTs), supports deployments, handles incidents, and maintains monitoring, observability, and operational documentation\n
  • Maintain all environments in an operational state across VAEC, including application/system hardware 24x7x365.\n
  • Apply OS and application patches and perform backup and restoration activities. \n
  • Support software releases (often outside core hours) and validate maximum load after production deployment. \n
  • Manage and optimize cloud resource utilization, including capacity planning and forecasting\n
  • Implement SRE practices and patterns to improve reliability, performance, automation, fault tolerance, telemetry, logging, and alerting across distributed and microservices architectures. \n
  • Improve fault tolerance for distributed/microservices systems when components or resources become temporarily unavailable. \n
  • Conduct trend analysis of system performance and resource utilization to forecast future demand and identify bottlenecks. \n
  • Implement realtime proactive monitoring dashboards enabling visibility into system health and alerting. \n, n \nPreferred:\n \n \n \n

  • Experience with AWS VA Enterprise Cloud (VAEC)\n
  • Experience with Azure DevOps\n

\n \nCompany Overview:\n Integral partners with federal defense, intelligence, and civilian leaders to tackle their most important challenges and deliver positive outcomes. Since our founding in 1998, we have helped clients leverage existing and emerging technologies to transform their enterprises, empower growth, drive innovation, and build sustainable success. The forward-leaning solutions we deliver are tailored to each mission with a focus on keeping our nation safe and secure.\n \n Integral is headquartered in McLean, VA and serves clients throughout the country.\n \n We offer a comprehensive total rewards package including paid parental leave and immediate vesting in our 401(k). Give us a try and become part of a curated group of professionals at Integral Federal!\n \n Our package also includes:\n \u2022 Medical, Dental & Vision Insurance\n \u2022 Flexible Spending Accounts\n \u2022 Short-Term and Long-Term Disability Insurance\n \u2022 Life Insurance\n \u2022 Paid Time Off & Holidays\n \u2022 Earned Bonuses & Awards\n \u2022 Professional Training Reimbursement\n \u2022 Employee Assistance Program\n \n Equal Opportunity Employer/Protected Veteran/Disability\n \n PI286561788”, “hiringOrganization”: {“@type”: “Organization”, “name”: “Integral Consulting Services”}, “jobLocation”: {“address”: {“addressCountry”: “United States”, “streetAddress”: “Not specified”, “@type”: “PostalAddress”, “postalCode”: “22102”, “addressLocality”: “Tysons”, “addressRegion”: “Virginia - VA”}, “@type”: “Place”}, “industry”: “”, “identifier”: {“@type”: “PropertyValue”, “name”: “Integral Consulting Services”, “value”: “286561788”}, “baseSalary”: {“@type”: “MonetaryAmount”, “currency”: “USD”, “value”: {“@type”: “QuantitativeValue”, “value”: “Competitive”, “unitText”

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.military.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:05 min

Practical Byzantine Fault Tolerance in distributed computing systems

Jonan Scheffler · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all