> Markdown version of [/jobs/ext/1421077-vnoc-resiliency-specialist](https://www.wearedevelopers.com/jobs/ext/1421077-vnoc-resiliency-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # VNOC Resiliency Specialist - **Company:** SteelGate LLC - **Location:** Alexandria, VA, United States (Remote available) - **Experience:** Starter - **Salary:** $80,000.0 - $85,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Performance Management, Systems Engineering, JIRA, Cloud Computing, Information Technology Operations, Queue Management Systems, Cloud Services, Runbook, Software Deployment, Data Logging, SC Clearance, SolarWinds (Software), Oracle Cloud Infrastructure, Servicenow - **Published:** July 24, 2026 - **Apply:** https://www.steelgatellc.com/job/systems-developer-journeyman/?form=apply#wpjb-scroll ## About the Role A successful candidate is a disciplined, detail-oriented operator who excels at executing established procedures. * Process-Oriented & Composed: You respect established SOPs and do not panic when critical alerts trigger. You can clearly follow step-by-step triage checklists during high-visibility incidents. * Autonomous & Reliable: You are self-motivated and punctual, capable of maintaining focus during quiet night-shift hours without direct leadership supervision. * Strong Communicator: You possess excellent written communication skills, essential for drafting clear chronological turnover logs and logging precise ticket updates. Target Experience: * 1 to 3 years of experience in an IT Operations Center (NOC/SOC) environment, Tier 1/2 IT Helpdesk, or relevant Military Operations (such as communications, cyber, or logistics). * Basic Cloud Knowledge: Basic understanding of cloud infrastructure concepts and exposure to AWS or OCI environments ## Description The VNOC Resiliency Specialist (Night Shift) supports 24/7 operations as the primary monitoring, routing, and triage authority alongside Incident Managers. Operating with high autonomy during off-hours, this team member serves as our first line of defense in maintaining system stability. Shift Information: The team uses a Pitman schedule (i.e., 2 days on, 2 days off, 3 days on, etc.) for 12-hour shifts to provide 24/7 coverage. The "typical" start-end time for the shift is 6:30p to 6:30a PT, but there can be some flexibility for the right candidate (e.g., a preference to start at 7:00p or 8:00p can likely be accommodated.) Location: This position is open to remote candidates, but candidates with proximity to Gilbert, AZ will be prioritized. Proximity to Gilbert makes it easier for the practitioner to attend trainings, meet with the broader team, participate in culture events, etc. Our Ideal Candidate: We are seeking a reliable, process-driven professional who remains composed under pressure and can thrive in an overnight environment. This role does not require deep system engineering; candidates with an eager attitude, aptitude, and ambition to grow are invited to apply. Core Functional Responsibilities * Dashboard Surveillance: Maintain monitoring of enterprise network, on-premise, and cloud infrastructure health utilizing SolarWinds, Elastic, and Application Performance Monitoring (APM) dashboards to catch system degradation or outages. * Queue Management & Triage: Triage incoming outage calls, acknowledge automated system alerts, prioritize the queue, and route events to the correct technical engineering teams based on established routing rules. * First-Response & Escalation: Provide composed, immediate first-line response during Major Incidents by executing standard step-by-step runbooks. Coordinate closely with Incident Managers to escalate issues and engage technical Subject Matter Experts (SMEs) as required. * Cloud Health Monitoring: Monitor high-level system alerts and service health dashboards within AWS and OCI (Oracle Cloud Infrastructure) environments to identify cloud service disruptions and initiate standard escalation workflows. * Ticketing Management: Manage the administrative lifecycle of incidents within ServiceNow and Jira, ensuring precise documentation of event timelines and ticket updates. * Shift Turnover & Incident Logging: Maintain precise shift turnover logs and conduct formal, detailed handovers to incoming day-shift personnel. Assist Incident Managers by compiling chronological event timeline data for post-incident reviews and After Action Reports (AARs). * Change Window Monitoring: Support change management activities by observing dashboard health statuses via SolarWinds and APM for anomalies during scheduled maintenance and software deployment windows tracked in ServiceNow. * Runbook & Process Adherence: Follow established VNOC runbooks, Tactical Techniques and Procedures (TTPs), and SOPs. Flag outdated documentation in Jira/ServiceNow repositories to ensure instructions remain accurate. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How to Write a CV and Interview if You Don't Fully Qualify For The Job](https://www.wearedevelopers.com/magazine/183-how-to-write-a-cv-and-interview-if-you-don-t-fully-qualify-for-the-job) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)