NOC Engineer / NOC Lead

STN, inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Shift work
Job source

Tech stack

Bash Shell Cloud Computing Linux Monitoring of Systems Issue Tracking Systems Information Technology Operations Python (Programming Language) Networking Basics Prometheus Runbook Datadog Scripting
+3 more
Grafana Performance Monitor Pagerduty

Job description

The NOC Engineer operates STN’s 24/7 monitoring and first-response capability for GPU One (GPUaaS) infrastructure. The role triages alerts, executes documented runbooks, and coordinates with on-call specialists during incidents to protect customer SLAs., * Monitor infrastructure alerts, customer SLA dashboards, and system health on a 24/7 basis

  • Triage incidents and engage on-call SREs, Network, Hardware, or Field Engineering as needed
  • Execute documented runbooks for common platform, network, and hardware issues
  • Manage the incident lifecycle including initial customer notification and status updates
  • Coordinate planned maintenance windows and change windows with internal teams and customers
  • Update status pages and customer-facing communications during incidents
  • Maintain shift handoff documentation and active-incident logs
  • Support ticket queue handling including Tier 1 ticket resolution
  • Contribute to continuous improvement of monitoring coverage, alert quality, and runbooks
  • Work rotating shifts including nights, weekends, and holidays

Requirements

Do you have experience in System performance monitoring?, * 3+ years in a NOC, SOC, or IT operations function

  • Hands-on experience with monitoring tools (Datadog, Prometheus, Grafana, PagerDuty, or equivalent)
  • Strong Linux and basic networking fundamentals
  • Excellent written and verbal communication, particularly under pressure
  • Willingness and ability to work rotating shifts including overnight coverage, * GPU, HPC, or large-scale cloud infrastructure background
  • ITIL Foundations certification
  • Demonstrated on-call and major-incident response experience
  • Scripting skills (Python, Bash) for runbook automation

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

9:06 min

Questions on career paths and continuous delivery orchestration platforms

Zan Markan Zan Markan · LIVE

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

3:08 min

Shifting software delivery bottlenecks to operations and incident response

Milin Desai Milin Desai +1 · WWC Europe 2026

Videos

See all

Related articles

See all