Network Operations Specialist - 2nd Shift
TQL
Cincinnati, United States of America
yesterday
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
IntermediateJob location
Cincinnati, United States of America
Tech stack
Microsoft Active Directory
Amazon Web Services (AWS)
Azure
Bash
Continuous Delivery
Continuous Integration
DevOps
Github
Monitoring of Systems
JSON
Python
Key Management
Network Troubleshooting
Windows Server
Network Service
Packet Analyzer
Citrix Systems
Operational Data Store
Powershell
Reliability Engineering
Runbook
Software Engineering
Systems Integration
Wide Area Networks
Datadog
Scripting (Bash/Python/Go/Ruby)
Google Cloud Platform
Computer Network Operations
SUSE Linux
System Availability
Reliability of Systems
Gitlab
GIT
Kubernetes
SolarWinds (Software)
REST
Splunk
New Relic (SaaS)
Webhooks
Dynatrace
Cisco networks
Docker
VMware
Programming Languages
Job description
The NOC Specialist is responsible for monitoring critical technology services, coordinating incident response, analyzing operational data, and supporting service restoration to maintain business continuity. This role helps develop and execute technical response plans, identify opportunities for proactive improvement, and reduce the impact of technology incidents on business operations.
The ideal candidate has experience working in a large enterprise environment and an interest in using automation, scripting, and DevOps practices to improve operational efficiency, monitoring, and incident response.
What You'll Be Doing:
- Provide incident management, event management, and operational monitoring in a 24x7x365 environment.
- Coordinate service restoration activities and identify opportunities for proactive service improvements.
- Support Senior NOC Specialists and Site Reliability Engineers with second-level network, server, application, and infrastructure incident management, with a focus on rapid service restoration.
- Monitor infrastructure, applications, network services, and operational alerts to identify potential service-impacting conditions.
- Participate in and lead incident bridges, communicate technical findings, and engage the appropriate support resources when incidents extend beyond the NOC's scope.
- Perform initial troubleshooting of server, network, application, and infrastructure-related incidents.
- Monitor and analyze system performance, availability, resource utilization, and capacity trends.
- Assist in maintaining 99.9% availability for critical business technology systems.
- Create and maintain operational documentation, troubleshooting procedures, technical manuals, and incident response plans.
- Work within established change management processes to implement approved system, monitoring, and operational updates.
- Mentor Associate NOC Specialists and help desk technicians to support their technical development and career progression.
- Plan and lead incident response exercises and training simulations to maintain companywide readiness.
- Identify repetitive operational activities that can be automated or standardized.
- Assist with the development and maintenance of scripts, monitoring integrations, dashboards, runbooks, and automated response workflows.
- Collaborate with infrastructure, application support, engineering, security, and DevOps teams to improve monitoring coverage, system reliability, and operational support processes.
- Participate in post-incident reviews and help identify corrective actions, monitoring improvements, and automation opportunities.
Requirements
- Ability to work M to F 3p to 12a
- Two to four years of relevant experience in network operations, infrastructure support, systems administration, application support, or a related technical field.
- Experience with incident management, event management, technical troubleshooting, or service restoration.
- Experience creating and maintaining technical documentation, operational procedures, and incident response plans.
- Foundational experience troubleshooting network, server, application, or infrastructure-related issues.
- Experience working in a large enterprise environment or supporting large-scale technology projects.
- Strong task-management, communication, and team-collaboration skills.
- Ability to communicate technical issues clearly during high-impact or time-sensitive incidents.
- Working knowledge of monitoring, alerting, system availability, and capacity-management concepts.
- Experience with technologies such as:
- Windows Server 2016/2019/2022
- Microsoft Active Directory
- VMware
- SUSE Linux
- Citrix
- Cisco routing and switching
- Packet captures and network troubleshooting
- General understanding of fiber, SD-WAN, and other enterprise WAN technologies.
Preferred Qualifications
- Experience using PowerShell, Python, Bash, or another scripting or programming language.
- Experience automating repetitive operational, monitoring, or incident-response tasks.
- Familiarity with DevOps, Site Reliability Engineering, infrastructure-as-code, or continuous integration and continuous delivery practices.
- Experience using source-control platforms such as Git, GitHub, GitLab, or Azure DevOps.
- Experience with monitoring and observability platforms such as Datadog, SolarWinds, Splunk, Dynatrace, New Relic, or similar tools.
- Familiarity with REST APIs, JSON, webhooks, and systems integrations.
- Experience creating automated runbooks, remediation workflows, dashboards, or alert-enrichment processes.
- Exposure to cloud platforms such as Microsoft Azure, Amazon Web Services, or Google Cloud.
- Familiarity with containers and orchestration technologies such as Docker or Kubernetes.
- Experience collaborating with application development, infrastructure engineering, automation, or DevOps teams.