Staff Network Engineer (Ai Fabric, Datacenter And Edge Networking) - Radian Arc (Emea)

Submer
A Coruña, Spain
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Remote
A Coruña, Spain

Tech stack

API
Artificial Intelligence
Computing Platforms
Bash
Border Gateway Protocol
Computer Clusters
Configuration Management
Computer Security
Computer Networks
Network Congestion
Wavelength-Division Multiplexing
Data Centers
DDoS Mitigation
Software Debugging
Distributed Computing Environment
Ethernet
Firmware
Monitoring of Systems
Network Topologies
Infrastructure as a Service (IaaS)
Networking Hardware
Python
Modular Design
Network Architecture
Network Control
Routing
Network Protocols
Citrix Systems
Performance Tuning
Remote Direct Memory Access
TensorFlow
Software Engineering
Virtual Local Area Networks
Computer Networking Systems
Cloud Platform System
PyTorch
Infrastructure Automation Frameworks
Low Latency
Network Optimization
Citrix Netscaler

Job description

Location & work modality:Europe/ RemoteStart:Aug Type of Contract:Full time or ContractAbout Radian ArcRadian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.What Impact You Will HaveDesign, implement, and operate the network infrastructure powering the GPU cloud platform, including high-performance AI fabrics as well as classical datacenter networking components such as routing, security, and external connectivity. This role spans both high-performance east-west networking for distributed AI workloads and north-south connectivity, security, and inter-datacenter transport.As the first dedicated networking role in the organization, the Staff Network Engineer combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands?on execution across design, deployment, troubleshooting, automation, and operational improvement.The Staff Network Engineer owns the long-term technical direction and operational strategy for Radian Arc's AI interconnect networks, designing scalable GPU fabrics and ensuring predictable low-latency performance across distributed training and inference workloads. The role includes designing large-scale RoCE and Ethernet fabrics, guiding architecture decisions, and ensuring operational excellence across global deployments, from hyperscale datacenters to smaller edge locations.You will collaborate closely with platform, compute, storage, observability, and operations teams to ensure networking is deeply integrated into the overall infrastructure architecture. This role also acts as the senior escalation point for complex networking incidents, driving deep technical investigations and systemic improvements that increase reliability, latency consistency, and operational maturity across the platform.Because this is currently the primary networking role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of technical direction, standards, cross-team influence, and long-term design, while also directly executing critical networking work that, in a larger organization, would be distributed across multiple engineers.What You'll DoAI Fabric & HPC NetworkingDesign and operate high-performance GPU networking fabrics supporting distributed AI workloads.Architect large-scale RoCE fabrics optimized for distributed training and inference.Optimize network performance for GPU communication patterns and east-west traffic.Design fabric topologies such as:Leaf-spineFat-TreeRail architecturesMulti-planeImplement high-performance networking technologies including:RDMARoCEHigh-bandwidth east-west fabricsSpectrum-XCollaborate with compute teams to support distributed training frameworks and GPU communication libraries.Define reference architectures and design principles for AI fabrics so future deployments follow reusable standards rather than one-off implementations.Evaluate architectural trade-offs across performance, resilience, cost, operability, and deployment speed, and make clear recommendations to stakeholders.Datacenter NetworkingDesign and operate Layer?2 and Layer?3 datacenter networks.Implement scalable routing architectures based on BGP and ECMP.Design tenant network isolation mechanisms across multi?tenant environments.Implement and maintain:Network bridgesRouting stacksOverlay networking systemsMaintain north?south ingress/egress routing and traffic management.Define standards and reusable patterns for segmentation, routing, and overlay integration across platform deployments.Security & Edge ConnectivityDeploy and maintain north?south security infrastructureImplement WAF and application?layer protectionsIntegrate security controls with platform services.Technologies Include:Citrix NetScaler / Citrix WAFTLS terminationDDoS mitigationAPI and proxy gateway protectionInter?Datacenter NetworkingDesign and operate private interconnects between datacentersImplement and maintain dark fiber ring architecturesOperate high-capacity WAN connectivity between regionsIntegrate datacenter fabrics into a global backbone networkDefine scalable design principles for backbone evolution, inter?site routing, redundancy, and failure-domain isolation.Technologies Include:DWDM / dark fiber transportBGP inter?site routingRedundant fiber ring architectures***G optical transportSpectrum-XGSEngineering Execution & DeliveryLead end-to-end engineering delivery of networking infrastructure, from design and labvalidation to production deploymentValidate network BOMs together with procurement and deployment teamsProvide detailed input into datacenter layouts and rack elevationsDrive capacity planning, performance modeling, and scaling strategiesEnsure network changes are executed safely with minimal customer impactAct as both the architectural owner and the practical execution lead for critical network initiatives during the build?out phase of the networking functionEstablish deployment standards, validation criteria, rollback approaches, and acceptance patterns that future engineers and teams can reuse.Operational Excellence & ReliabilityOwn operational performance and reliability of networking infrastructureDrive automation for:ProvisioningConfiguration managementMonitoringLifecycle managementImprove day?2 operations through automation and operational toolingLead incident response and root?cause analysis for major network eventsDefine and track SLAs, SLOs, and reliability metricsTranslate major incidents and operational pain points into durable standards, design changes, and long?term architectural improvementsEstablish measurable benchmarks for reliability, latency consistency, operability, and recovery behavior across network deployments.Cross?Functional CollaborationWork closely with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams.Provide technical leadership across infrastructure initiatives.Communicate architectural decisions, trade?offs, and risks clearly to stakeholders.Influence the long?term platform networking roadmap and architecture.Act as the primary networking design authority across the organization, guiding adjacent teams on how networking constraints and capabilities should shape platform decisions.Raise the technical bar by mentoring engineers in adjacent domains and helping build the future networking function.Technical StackDatacenter NetworkingBGPEVPN / VXLANECMPVLAN / VRFOVS / OVNLinux networkingBlueField DPURouting & Control PlaneVyOSBGP-based routing architecturesECMP fabricsSecurityCitrix NetScaler / WAFDDoS protectionTransport & BackboneDark fiberMetro fiber ringsDWDM transport***G optical networkingAI NetworkingRDMARoCEGPU fabricsLarge-scale east?west compute networkingCongestion controlWhat You'll NeedCore ExperienceStrong hands?on experience designing and operating large-scale datacenter networksExpert knowledge of modern networking protocols including:BGPOSPFECMPEVPN / VXLANProven experience operating high-speed Ethernet networks in production environmentsExperience operating NVIDIA / Mellanox networking platformsExperience owning both architecture and direct implementation in lean or fast?scaling environments is strongly preferred.Advanced AI Fabric Networking ExpertiseThe candidate should have deep expertise in designing and operating networking fabrics optimized for large-scale GPU clusters and distributed AI workloads. This includes a strong understanding of GPU communication patterns and the networking requirements of distributed training and inference systems.Deep understanding of NCCL communication patterns and their impact on network topology and performance.Experience tuning RoCE fabrics for large-scale GPU clusters.Strong knowledge of RDMA transport behavior and failure modes.Practical experience implementing and tuning PFC and ECN for congestion management.Understanding of GPU collective communication patterns such as all?reduce, all?gather, broadcast, reduce?scatter, and their impact on east?west network traffic.Experience designing rail?optimized GPU networking fabrics for distributed training and inference clusters.Familiarity with diagnosing performance issues related to:NCCL stallsRDMA congestionFabric hotspotsPacket loss impacting distributed trainingUnderstanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow.The candidate should also be able to collaborate closely with compute platform teams to ensure that networking infrastructure is optimized for distributed training, distributed inference, and high-throughput AI workloads.Systems & TroubleshootingAbility to debug complex cross-layer issues spanning:HardwareFirmwareKernel networkingDistributed application communication layersStrong knowledge of networking hardware, optics, and high-speed interconnects.Experience designing network observability systems.Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues.AutomationStrong automation skills using Python and/or Bash.Experience applying software engineering practices to infrastructure automation.Experience building reusable tooling, standards, or validation approaches that increase leverage across teams.LeadershipProven ability to lead complex technical initiatives across teams.Comfortable collaborating across engineering, operations, and vendors.Strong systems?level thinking balancing performance, reliability, scalability, and operational cost.Demonstrated ability to set architectural direction and drive adoption of engineering standards across an organization.Proven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority.Strong mentoring capability and ability to raise the technical level of adjacent engineering teams.Able to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency.What We OfferAttractive compensation package reflecting your expertise and experience.A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid?friendly approach.You'll be part of a fast?growing scale?up with a mission to make a positive impact, offering an exciting career evolution.Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.Our inclusive responsibilityRadian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.#J--Ljbffr

Requirements

Proven experience operating high-speed Ethernet networks in production environments Experience operating NVIDIA / Mellanox networking platforms Experience owning both architecture and direct implementation in lean or fast?scaling environments is strongly preferred. Advanced AI Fabric Networking Expertise The candidate should have deep expertise in designing and operating networking fabrics optimized for large-scale GPU clusters and distributed AI workloads. This includes a strong understanding of GPU communication patterns and the networking requirements of distributed training and inference systems. Deep understanding of NCCL communication patterns and their impact on network topology and performance. Experience tuning RoCE fabrics for large-scale GPU clusters. Strong knowledge of RDMA transport behavior and failure modes. Practical experience implementing and tuning PFC and ECN for congestion management. Understanding of GPU collective communication patterns such as all?reduce, all?gather, broadcast, reduce?scatter, and their impact on east?west network traffic. Experience designing rail?optimized GPU networking fabrics for distributed training and inference clusters. Familiarity with diagnosing performance issues related to: NCCL stalls RDMA congestion Fabric hotspots Packet loss impacting distributed training Understanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow. The candidate should also be able to collaborate closely with compute platform teams to ensure that networking infrastructure is optimized for distributed training, distributed inference, and high-throughput AI workloads. Systems & Troubleshooting Ability to debug complex cross-layer issues spanning: Hardware Firmware Kernel networking Distributed application communication layers Strong knowledge of networking hardware, optics, and high-speed interconnects. Experience designing network observability systems. Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues. Automation Strong automation skills using Python and/or Bash. Experience applying software engineering practices to infrastructure automation. Experience building reusable tooling, standards, or validation approaches that increase leverage across teams. Leadership Proven ability to lead complex technical initiatives across teams. Comfortable collaborating across engineering, operations, and vendors. Strong systems?level thinking balancing performance, reliability, scalability, and operational cost. Demonstrated ability to set architectural direction and drive adoption of engineering standards across an organization. Proven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority. Strong mentoring capability and ability to raise the technical level of adjacent engineering teams. Able to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency.

Benefits & conditions

Attractive compensation package reflecting your expertise and experience. A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid?friendly approach. You'll be part of a fast?growing scale?up with a mission to make a positive impact, offering an exciting career evolution. Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands. Our inclusive responsibility Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law. #J-*****-Ljbffr

About the company

About Radian Arc Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments. What Impact You Will Have Design, implement, and operate the network infrastructure powering the GPU cloud platform, including high-performance AI fabrics as well as classical datacenter networking components such as routing, security, and external connectivity. This role spans both high-performance east-west networking for distributed AI workloads and north-south connectivity, security, and inter-datacenter transport.

Apply for this position