NOC Engineer III, Incident Management Team

Google LLC
Austin, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$134,000.0 - $152,000.0
Working hours
Shift work
Job source

Tech stack

Border Gateway Protocol Wavelength-Division Multiplexing Dynamic Host Configuration Protocol Issue Tracking Systems Networking Hardware Network Monitoring Open Shortest Path First (OSPF) Network Routing Diagnostic Tools Mttr Juniper Information Technology
+1 more
ArcSight Event Correlation

Job description

As an Incident Manager, you’ll perform ticket administration, event correlation, diagnostics, issue repair, ensuring incidents are prioritized based on their business impact and customer satisfaction. As a Single Thread Owner, you will respond to all reported incidents initiating the proper management process with the appropriate teams to restore service as quickly as possible by developing and implementing incident resolution plans. You’ll have the opportunity to interact with internal stakeholders across different shifts and teams, sharing regular reports on incident metrics, post mortems, industry updates, influencing efficiencies and more.

In this role, you’ll:

  • Lead the end-to-end remediation of all high severity events (critical escalation paths, compliance with on-call duties, vendor rolodex, augmentation of existing support models from vendors, and post mortem management).
  • Initiate and lead war room phone calls with cross-functional stakeholders, including but not limited to: network engineers, system administrators, customer support, and management to drive resolution.
  • Delegate tasks, track progress, and hold stakeholders accountable for timely completion of assigned actions.
  • Run stakeholder comms and deliver written summaries and reports to executive leadership teams about the incident status, progress, and estimated time to resolution (ETR).
  • Assist in the generation of automation. Identify, triage, and implement preventive measures to reduce the frequency and severity of incidents and all “chronic/repetitive” issues, assisting in the generation of automation.
  • Generate post mortem/RCA, identifying lessons learned, actions and drive to fruition, and regular reports on incident metrics, including response times, resolution rates and KPIs to track improvements in MTTR.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, telecommunications, a related field, or equivalent practical experience.
  • 3 years of experience with network routing protocols, design and troubleshooting, with network equipment providers.
  • Experience with network monitoring, ticketing systems, triaging escalation tools, and troubleshooting tools.
  • Experience with TCP/IP networking concepts, including BGP and OSPF routing protocols, switching and DHCP.
  • Ability to work non-standard working hours including nights, weekends, holidays, and differing work rotations/shifts.

It’s preferred if you have:

  • Experience in incident management, including leading war room calls, driving communication with stakeholders, preparing executive summaries, and providing timely updates and estimated time to resolution (ETR) during network outages or critical incidents.
  • Technical certifications, such as Nokia (ONC or NRS), Ciena(CE-A or CE-P), MEF CECP, Juniper (JNCIA or JNCIP), FOT, FTAA, DCI, WNT or MRT.
  • Experience in Radio Engineering incident management and switched network implementations.
  • OSS Functionality (Fault Management and ticket administration).
  • Understanding of CWDM/DWDM theory (C-band, L-band, wavelengths), linear and ring topologies, network hierarchy, optical and routed Networks, and ITIL knowledge.
  • Experience with the following equipment: CWDM/DWDM; Ciena Waveserver; Juniper MX/QFX/PTX, Nokia 7x50, Adtran/Nokia PON/OLT/ONT.

Benefits & conditions

The US base salary for this full-time position is between $134,000 - $152,000 + bonus + benefits. As pay varies by location, your recruiter will share more about the specific salary range for your targeted location during the hiring process.

About the company

At GFiber, we believe that great internet has the power to drive innovation, strengthen communities, enable the impossible, and do all the everyday things that make all of our world go round. And the job of creating better internet is never done - so we’re growing! Our team is committed to building a place where people who want to make a difference can grow their careers and find their spot to belong.

GFiber is an Alphabet company that brings Google Fiber and Google Fiber Webpass internet services to homes and businesses across the United States. Our teams are expanding as we connect more cities and people to exceptional internet., GFiber’s mission is to deliver abundant internet on networks that are always fast and always open with products that are easy to understand and clearly priced. We believe customers deserve a better internet experience and everything we do is focused on providing just that. On our team, you’ll work in an environment that’s redefining the status quo in the Internet industry. On GFiber’s Incident Management team, our NOC engineers resolve outages efficiently with their knowledge of network protocols, configurations, and troubleshooting techniques. This translates to shorter periods of downtime, preventing outages altogether, or identifying them in their early stages, minimizing downtime and improving customer experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 ¡ Coffee With Developers

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley ¡ WWC 2021

4:43 min

Connecting multiple container namespaces using virtual network bridges

Oliver Seitz Oliver Seitz ¡ WWC 2025

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz ¡ LIVE

2:29 min

Synchronizing global routing tables permissionlessly across miner networks

Lin Zheming ¡ LIVE

4:19 min

Introduction to network security and endpoint monitoring architectures

Christoph Ruggenthaler ¡ LIVE

Videos

See all

Related articles

See all