Cloud SRE Systems Engineer

Integral Consulting Services
Tysons, VA, United States
13 days ago
Apply on www.jofdav.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Systems Engineering JIRA Microsoft Azure Cloud Computing Fault Tolerance Github Reliability Engineering Site Reliability Engineering Practices Software Deployment Backup and Restore Data Logging HybridCloud
+3 more
Information Technology Performance Monitor Microservices

Job description

The Cloud Site Reliability Engineering (SRE) Systems Engineer is responsible for ensuring reliability, availability, performance, and operational excellence of VA product environments deployed in the VA Enterprise Cloud (VAEC) for the Department of Veterans Affairs (VA), Office of Information and Technology (OIT), Product Delivery Services (PDS), Benefits and Memorials (BAM), Memorial and Benefit Services (MBS) deliver secure, reliable, and effective Information Technology (IT) solutions that support the Department’s mission., * Applies SRE practices, manages cloud environments, performs patching and automation, ensures compliance with required Service Level Targets (SLTs), supports deployments, handles incidents, and maintains monitoring, observability, and operational documentation

  • Maintain all environments in an operational state across VAEC, including application/system hardware 24x7x365.
  • Apply OS and application patches and perform backup and restoration activities.
  • Support software releases (often outside core hours) and validate maximum load after production deployment.
  • Manage and optimize cloud resource utilization, including capacity planning and forecasting
  • Implement SRE practices and patterns to improve reliability, performance, automation, fault tolerance, telemetry, logging, and alerting across distributed and microservices architectures.
  • Improve fault tolerance for distributed/microservices systems when components or resources become temporarily unavailable.
  • Conduct trend analysis of system performance and resource utilization to forecast future demand and identify bottlenecks.
  • Implement realtime proactive monitoring dashboards enabling visibility into system health and alerting.

Requirements

  • Bachelor’s Degree in computer science or IT related degree with 7-10 years’ experience
  • Knowledge of modern SRE practices: automation, resilience engineering, observability, fault tolerance, SLT/SLA adherence, and incident analysis.
  • Experience producing monitoring reports, SLT assessments, postmortems, and operational dashboards
  • Experience with Jira and GitHub
  • Competency maintaining IRPs, DRPs, backup/restore plans, and executing continuity operations
  • Strong documentation and communication skills for reporting outages, SLT performance, and incident updates
  • Public Trust

Preferred:

  • Experience with AWS VA Enterprise Cloud (VAEC)
  • Experience with Azure DevOps

Benefits & conditions

We offer a comprehensive total rewards package including paid parental leave and immediate vesting in our 401(k). Give us a try and become part of a curated group of professionals at Integral Federal!

Our package also includes:

· Medical, Dental & Vision Insurance

· Flexible Spending Accounts

· Short-Term and Long-Term Disability Insurance

· Life Insurance

· Paid Time Off & Holidays

· Earned Bonuses & Awards

· Professional Training Reimbursement

· Employee Assistance Program

About the company

Integral partners with federal defense, intelligence, and civilian leaders to tackle their most important challenges and deliver positive outcomes. Since our founding in 1998, we have helped clients leverage existing and emerging technologies to transform their enterprises, empower growth, drive innovation, and build sustainable success. The forward-leaning solutions we deliver are tailored to each mission with a focus on keeping our nation safe and secure.

Integral is headquartered in McLean, VA and serves clients throughout the country.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

3:46 min

Navigating a career in cloud transformation consulting

Piet Van Dongen · LIVE

Videos

See all

Related articles

See all