> Markdown version of [/jobs/ext/2035484-group-lead-infrastructure-services](https://www.wearedevelopers.com/jobs/ext/2035484-group-lead-infrastructure-services). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Group Lead - Infrastructure Services - **Company:** DriveNets Ltd. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Cloud Engineering, Computer Clusters, Computer Engineering, System Configuration, Data Centers, Linux, Ethernet, InfiniBand, Routing, Performance Tuning, Remote Direct Memory Access, Broadcom, Tensorflow, Prometheus, Google Cloud, Pytorch, Large Language Models, Grafana, Data Center Networking, Kubernetes, Information Technology, Slurm, Oracle Cloud Infrastructure, Cerner CCL - **Published:** August 12, 2026 - **Apply:** https://drivenets.com/job/?id=66.D61 ## About the Role What we need to see: * 10+ years of experience in AI/HPC infrastructure, data center networking, or solutions architecture, with at least 2-3 years in a technical leadership or team lead capacity. * Hands-on technical depth across both compute infrastructure (GPU clusters, Linux systems, AI workloads) and data center networking (routing, switching, fabric design), with the ability to engage credibly across both disciplines. * Proven experience leading customer-facing technical teams through complex deployment and POC cycles in AI/HPC or data center environments. * Strong understanding of AI cluster architecture - including GPU platforms (NVIDIA, AMD), RDMA networking, storage connectivity, and the interaction between compute, network, and storage layers. * Experience with performance benchmarking methodologies (NCCL/RCCL, RDMA, LLM workloads) and the ability to interpret and act on results at a system level. * Demonstrated ability to work cross-functionally with Sales, Product Management, and Engineering teams, translating customer feedback into product improvements and go-to-market strategy. * Excellent communication and presentation skills, with proven ability to influence technical and executive stakeholders at customer organizations. * Ability to write extensive technical content (white papers, technical briefs, design guides, etc.) for external audiences with a balance of technical accuracy and clear messaging. * Ability to travel domestic and international. Ways to stand out from the crowd: * Deep familiarity with AI-relevant infrastructure technologies - InfiniBand, RoCEv2, lossless Ethernet (PFC, ECN), GPU, NIC, DPU, and accelerated computing platforms. * Hands-on experience deploying and operating large-scale AI/HPC clusters, including GPU resource scheduling (Slurm, Kubernetes), monitoring (Prometheus, Grafana, DCGM), and operational tooling. * Understanding of scale-up (NVLink, UALink) and scale-out (Enhanced Ethernet, UEC, InfiniBand) interconnect technologies and their design trade-offs. * Experience with CCL tuning (NCCL/RCCL), GPU environment setup, and performance optimization across large multi-node GPU clusters. * Familiarity with AI/ML frameworks (PyTorch, TensorFlow) and how workload characteristics interact with infrastructure design decisions. * Proven experience with one or more Tier-1 Clouds (AWS, Azure, GCP, or OCI) or emerging NeoClouds, and cloud-native architectures and software. * Background in data center operations fundamentals - networking, cooling, power, and rack-level design. * Experience engaging compute, NIC, or storage vendors on joint solution definition, reference architecture development, or benchmarking programs., BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or other Engineering fields, or equivalent experience. ## Description DriveNets is seeking a Group Lead for its Infrastructure Services (DIS) team to be a key member of our customer-facing technical organization. Join a dynamic and forward-thinking company at the forefront of AI infrastructure. We leverage advanced technologies to develop innovative solutions that drive efficiency, scalability, and exceptional compute performance. Collaborate with the industry's best as we partner with hyperscalers, emerging NeoClouds, and enterprises building large-scale AI/HPC clusters, shaping the future of disaggregated AI networking and compute infrastructure. Our environment fosters creativity, teamwork, and growth, and offers you the opportunity to make a meaningful impact while leading a high-performing team on groundbreaking deployments. As Group Lead for DIS, you will manage and develop a team of Solution Engineers and Solutions Architects responsible for designing, deploying, and optimizing DriveNets' AI/HPC infrastructure solutions at customer sites. You will provide technical leadership across the full customer lifecycle - from pre-sales architecture and POC execution through deployment, performance benchmarking, and ongoing operations. You will work cross-functionally with Sales, Product Management, and Engineering to ensure customer success, drive product feedback, and continuously raise the bar for technical delivery quality across the team. Responsibilities * Lead and develop the DIS team - a group of Solution Engineers and Solutions Architects - setting technical direction, managing execution, and fostering a culture of ownership, learning, and customer focus. * Oversee end-to-end customer engagement for DIS - from pre-sales technical support and solution architecture through POC planning, deployment execution, and post-deployment operations. * Serve as the senior technical escalation point for customer infrastructure challenges, including AI cluster performance issues, networking design trade-offs, and operational reliability concerns. * Partner with Sales Account Managers to support business opportunities, lead technical responses to RFP/RFQs, and influence technical decision-makers at the VP and CxO level. * Guide the team in conducting performance benchmarking activities - including NCCL/RCCL, RDMA, and LLM benchmarks - and ensure results are translated into actionable product and deployment insights. * Work with Product Management and Engineering to funnel customer requirements, field observations, and performance data into the product roadmap and development backlog. * Define and drive internal processes for deployment planning, operational readiness, monitoring standards, and technical documentation across the DIS team. * Build and maintain relationships with compute, NIC, and storage partners to support joint POCs, reference deployments, and solution validation. * Represent DriveNets at industry events and conferences, and contribute to external technical content including white papers, blogs, and design guides. * Recruit, mentor, and grow team members, and establish clear performance goals aligned with business objectives. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)