> Markdown version of [/jobs/ext/3583793-principal-network-developer](https://www.wearedevelopers.com/jobs/ext/3583793-principal-network-developer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Network Developer - **Company:** Oracle - **Location:** Nashville, TN, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Border Gateway Protocol, Cloud Computing, Complex Networks, Network Congestion, Data Centers, Distributed Systems, Ethernet, Python (Programming Language), Network Architecture, Routing, Open Shortest Path First (OSPF), Oracle (Applications), Remote Direct Memory Access, Reliability Engineering, Cloud Services, Ansible, Software Engineering, AI Infrastructure, Graphics Processing Unit (GPU), Computer Networking Systems, Computer Network Operations, High Performance Computing, Data Center Networking, Open Network Automation Platform, Oracle Cloud Infrastructure - **Published:** October 4, 2026 - **Apply:** https://eeho.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1/requisitions/preview/345549 ## About the Role * Experience designing, operating, and troubleshooting large-scale RDMA/RoCE, cloud, data center, or high-performance networks. * Knowledge of RDMA, RoCE, Ethernet fabrics, congestion control, QoS, and AI/GPU networking. * Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, and data center networking. * Experience with network automation using Python, Ansible, APIs, or similar technologies. * Experience with network telemetry, observability, monitoring, performance analysis, and incident management. * Experience supporting hyperscale cloud, AI/GPU, HPC, or large-scale distributed infrastructure. * Ability to lead complex technical initiatives and collaborate across engineering, operations, customers, and vendors. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. ## Description Join the AI Network Operations team to design, operate, and optimize advanced network systems supporting large-scale AI and cloud infrastructure. As an experienced Network Engineer, you will help operate high-performance RDMA network fabrics powering tier-0 customers in the generative AI industry, while partnering across engineering teams and vendors to ensure performance, reliability, and scalability. The OCI AI Infrastructure - Network Operations team operates the high-performance RDMA/RoCE network fabrics powering OCI's largest AI, GPU, and HPC workloads. These globally distributed networks support some of the most demanding generative AI workloads running on Oracle Cloud Infrastructure. As a Principal Network Engineer, you will lead the design, deployment, operation, and optimization of large-scale RDMA/RoCE network fabrics across OCI's global cloud infrastructure. You will combine deep networking expertise with strong automation and software engineering skills to improve network performance, scalability, reliability, and operational efficiency. You will design advanced automation, testing, telemetry, and monitoring solutions; lead network validation, incident response, and root cause analysis; and drive performance and capacity improvements across production environments. You will partner with engineering teams, vendors, and customers to resolve complex technical challenges, ensure deployment readiness, and evolve network architecture and operational practices at cloud scale. As a technical leader, you will also mentor engineers, influence architecture and engineering standards, and help shape the tools and systems supporting hundreds of thousands of network devices and millions of servers across OCI., * Lead the design, deployment, validation, and lifecycle management of large-scale RDMA/RoCE network fabrics supporting OCI AI, GPU, and HPC infrastructure. * Translate network architectures into scalable designs and deployment plans, ensuring performance, reliability, and operational readiness across OCI's global cloud environment. * Serve as technical lead for complex network initiatives spanning RDMA/RoCE fabrics, data center networking, automation, testing, deployment, and operations. * Develop automation frameworks, tools, scripts, and infrastructure pipelines to improve network deployment, testing, reliability, and operational efficiency. * Design test strategies and lead pre-production validation, network change reviews, and deployment readiness for high-performance network fabrics. * Build and enhance telemetry, monitoring, dashboards, and alerting to identify network health, congestion, performance, and reliability issues. * Lead incident response, complex troubleshooting, root cause analysis, and corrective actions for network issues impacting OCI AI and GPU workloads. * Analyze network performance and capacity, including latency, throughput, packet loss, and congestion, to drive scalable improvements. * Partner across OCI Network Engineering, SRE, AI Infrastructure, Data Center Operations, product teams, and vendors to deliver reliable network solutions. * Mentor engineers and contribute to network architecture, engineering standards, operational tooling, and continuous improvement.