> Markdown version of [/jobs/ext/2722810-unified-fabric-manager](https://www.wearedevelopers.com/jobs/ext/2722810-unified-fabric-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Unified Fabric Manager - **Company:** NVIDIA Ltd. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $170,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Artificial Intelligence, Systems Engineering, Border Gateway Protocol, Big Data, Common Lisp Object Systems, Computer Clusters, Computer Networks, Network Congestion, Data Centers, Firmware, InfiniBand, Python (Programming Language), Network Troubleshooting, Network Layer, Linux System Administration, Network Architecture, Network Monitoring, Routing, Remote Direct Memory Access, Ansible, Prometheus, AI Infrastructure, Graphics Processing Unit (GPU), System Availability, Grafana, Git, AI Platforms, Hardware Infrastructure, Restful APIs, Terraform, Open Network Automation Platform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-network-engineer-infiniband-ufm-lightning-ai-8776629 ## About the Role * 7+ years of experience in large-scale data center networking. * Experience deploying, operating, and troubleshooting InfiniBand environments. * Strong understanding of Layer 2 and Layer 3 networking, including BGP, EVPN, VXLAN, and spine-leaf architectures. * Strong Linux administration experience. * Experience with network automation using Python, Ansible, Terraform, or similar tools. * Experience with network troubleshooting, observability, and telemetry tooling. * Excellent troubleshooting, documentation, and communication skills., * Experience with Netris, Terraform, or similar infrastructure tooling. * Experience with multi-region backbone design. * Exposure to bare-metal provisioning systems. * Experience operating HPC, GPU-dense, or other large-scale infrastructure environments. * Experience working in high-growth infrastructure startups ## Description We are seeking an experienced Senior Network Engineer (InfiniBand / UFM) to design, deploy, automate, and operate next-generation AI Factory networking infrastructure supporting large-scale GPU clusters. This role is responsible for building and maintaining high-performance NVIDIA Quantum InfiniBand fabrics that power AI training and inference environments utilizing NVIDIA UFM, NCCL, RoCEv2, and modern data center technologies. The ideal candidate has deep expertise in InfiniBand networking, NVIDIA Unified Fabric Manager (UFM), large-scale GPU deployments, Linux networking, automation, and troubleshooting distributed AI workloads. The network is the foundation of distributed AI training. Performance bottlenecks at the fabric layer directly impact customer workloads. This role will shape how Lightning scales from tens of thousands of GPUs to significantly beyond - ensuring deterministic performance, predictable scaling, and enterprise-grade reliability. If you want to architect infrastructure that powers frontier AI research, this is the role for you. This role is hybrid with a minimum of 2 in-office days per week in San Francisco, Seattle, or NYC, with fully remote work considered for candidates outside of our office hub locations. All employees participate in occasional team and company offsites. What You'll Do * Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters. * Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health. * Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches. * Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs. * Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures. * Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure. * Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication. * Work closely with AI platform, GPU infrastructure, storage, and systems engineering teams to deploy scalable AI Factory environments. * Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies. * Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms. * Support high availability, maintenance windows, incident response, root cause analysis, and capacity planning. * Participate in architecture reviews and define networking standards for AI infrastructure. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)