Production Network Engineer

The Meta Game, Inc.
Providence, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 227K

Job location

Providence, United States of America

Tech stack

Artificial Intelligence
Border Gateway Protocol
Big Data
Network Operating System (NOS)
Network Analysis
Configuration Management
Complex Networks
Computer Engineering
Data Centers
Distributed Systems
Ethernet
Multi-protocol Systems
Python
Network Architecture
Network Planning and Design
Routing
Network Service
Network administration
Open Shortest Path First
Software Engineering
Scripting (Bash/Python/Go/Ruby)
Computer Network Operations
Computer Network Technologies
Reliability of Systems
Information Technology
SDN Network
Open Network Automation Platform
Programming Languages

Job description

Meta's global network infrastructure underpins billions of user connections and powers some of the world's most demanding AI and distributed computing workloads. The Production Network Engineering team is responsible for the design, deployment, operation, and continuous improvement of Meta's large-scale production network - spanning data centers, edge points of presence, and backbone interconnects. In this role, you will drive strategy and execution across complex network systems, lead cross-functional initiatives to improve reliability and performance, and apply deep subject matter expertise to solve problems at a scale few networks in the world match., 1. Lead the design and delivery of large-scale production network projects spanning data center fabrics, backbone infrastructure, and edge network services

  1. Define and drive the technical strategy and roadmap for network reliability, capacity, and operational efficiency across multiple teams
  2. Develop and implement automation frameworks to reduce manual operational overhead and accelerate network design synthesis, network build and configuration management
  3. Own cross-functional incident response and post-incident review processes, driving root cause analysis and systemic improvements to reduce recurrence
  4. Identify and resolve complex network performance, routing, and reliability issues across multi-vendor, multi-protocol production environments
  5. Collaborate with network architecture, capacity planning, and software engineering teams to align infrastructure investments with evolving AI and product demand forecasts
  6. Establish and refine operational standards, runbooks, and monitoring frameworks to improve network observability and reduce mean time to detection and resolution
  7. Contribute to organizational strategy by defining scalable approaches to network operations and influencing tooling and platform decisions across engineering teams
  8. Mentor other engineers on network engineering best practices, operational discipline, and systems thinking across the production environment
  9. Leverage AI-integrated workflows to accelerate network analysis, anomaly detection, and documentation, sharing learnings to scale adoption across the team

Requirements

  1. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  2. 8+ years of experience in production network engineering, including design, deployment, and operations of large-scale data center or backbone network infrastructure
  3. Experience with routing protocols and network technologies including BGP, OSPF, IS-IS, MPLS, and large-scale Ethernet fabrics in a production environment
  4. Experience developing and deploying network automation using scripting or programming languages such as Python to manage configuration, provisioning, or monitoring at scale
  5. Experience leading cross-functional network reliability or infrastructure projects, including driving incident response, root cause analysis, and systemic remediation
  6. Experience influencing technical decisions and network strategy across multiple engineering teams through written proposals, design reviews, and stakeholder alignment, 17. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  7. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  8. Experience operating and troubleshooting networks at hyperscale, including multi-vendor spine-leaf data center fabrics and global backbone interconnects
  9. Experience with network observability platforms, traffic engineering, and capacity modeling in high-throughput production environments
  10. Experience integrating AI tools to redesign network operations workflows and deliver measurable improvements in efficiency or reliability outcomes
  11. Familiarity with software-defined networking principles, network operating system internals, or co-development of network management platforms with software engineering teams
  12. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies

Apply for this position