> Markdown version of [/jobs/ext/1810027-network-engineer-ai-infrastructure-repair](https://www.wearedevelopers.com/jobs/ext/1810027-network-engineer-ai-infrastructure-repair). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Network Engineer, AI Infrastructure Repair - **Company:** The Meta Game, Inc. - **Location:** Little Rock, AR, United States - **Salary:** $193,000.0 - $271,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Complex Networks, Computer Engineering, Data Centers, Ethernet, InfiniBand, Networking Hardware, Machine Learning, Network Architecture, Remote Direct Memory Access, Data Driven Tests, Runbook, AI Infrastructure, Datadog, Computer Network Operations, High Performance Computing, Reliability of Systems, Information Technology - **Published:** July 7, 2026 - **Apply:** https://dejobs.org/x/x/7B710045667B4F86888FA8FCC20A14BF/job/ ## About the Role Network Engineer, AI Infrastructure Repair Responsibilities, 11. Experience influencing technical direction and organizational strategy through data-driven analysis, written proposals, and stakeholder alignment across engineering and operations teams 12. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 13. Experience leading cross-functional programs that span network operations, hardware deployment, and infrastructure reliability at data center scale 14. Experience developing and driving strategy for network fault management, repair automation, or remediation programs in production environments 15. Experience designing, deploying, or operating high-speed network fabrics used in AI or machine learning infrastructure, including technologies such as RDMA over Converged Ethernet, InfiniBand, or high-density optical interconnects 16. 12+ years of experience in network engineering, with a focus on large-scale data center or high-performance computing network environments, 17. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 18. Experience with network telemetry platforms, observability tooling, or AI-assisted anomaly detection applied to large-scale fabric environments 19. Experience building or scaling repair operations programs, including workforce planning, tooling development, and process standardization across multiple data center sites 20. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) 21. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) 22. Track record of contributing to network hardware or topology design reviews, translating operational repair insights into upstream engineering improvements 23. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 24. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 25. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 26. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 27. Familiarity with AI accelerator interconnect architectures and the network reliability requirements of distributed training workloads at hyperscale ## Description Meta is building the next generation of AI infrastructure to power large-scale machine learning workloads, and the reliability of that infrastructure depends on reliable, high-performance network engineering. In this role, you will lead the strategy and execution for AI network repair and remediation programs, ensuring that the high-performance fabrics underpinning Meta's AI training and inference clusters remain operational, resilient, and optimized. You will drive cross-functional initiatives spanning network deployment, fault diagnosis, and repair automation across Meta's AI data center environments, shaping the systems and processes that keep AI infrastructure at scale., 1. Define and drive the long-term strategy for AI network repair and remediation programs across large-scale data center environments supporting machine learning workloads 2. Lead root cause analysis and resolution of complex network faults affecting high-performance AI training and inference fabrics, including RDMA, high-speed Ethernet, and optical interconnect layers 3. Develop and champion novel approaches to network fault detection, automated remediation, and repair workflow optimization for AI cluster infrastructure 4. Partner with hardware, software, and data center operations teams to align network repair programs with AI infrastructure deployment roadmaps and capacity plans 5. Establish and refine operational frameworks, runbooks, and tooling for network repair at scale, reducing mean time to repair across AI fabric environments 6. Identify systemic reliability risks in AI network infrastructure and drive cross-functional initiatives to address them before they impact production workloads 7. Influence the design of next-generation AI network architectures by contributing repair and reliability insights to hardware and topology decisions 8. Leverage AI-driven analytics and automation tools to redesign repair workflows, accelerating fault identification and resolution across distributed network environments 9. Build and maintain strategic relationships with internal engineering, operations, and vendor partners to ensure repair programs scale with AI infrastructure growth 10. Communicate program status, risk, and strategic recommendations to engineering leaders and cross-functional stakeholders through structured reporting and executive briefings ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)