> Markdown version of [/jobs/ext/626605-software-engineer-ai-middleware](https://www.wearedevelopers.com/jobs/ext/626605-software-engineer-ai-middleware). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - AI Middleware - **Company:** Cornelis Networks - **Location:** Austin, TX, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Artificial Intelligence, C++ (Programming Language), Code Review, Nvidia CUDA, Linux, Distributed Computing Environment, Ethernet, Remote Direct Memory Access, Tensorflow, System Programming, Performance Testing, Pytorch, Free and Open-Source Software, Cerner CCL - **Published:** June 24, 2026 - **Apply:** https://www.dice.com/job-detail/0fc3e4df-2950-4ad8-95cd-ff8408a9eb54 ## About the Role * 8+ years of experience in high-performance systems programming in C/C++ on Linux. * Strong experience with GPU communication stacks including CUDA/ROCm and NCCL/RCCL. * Ability to optimize distributed training performance using profiling and tracing. * Understanding of collective communication concepts and topology awareness. * Experience delivering production-quality code. * Open-source contributions in relevant areas. Preferred Qualifications * Experience with AI frameworks such as PyTorch Distributed, DeepSpeed, and Megatron-LM. * Familiarity with libfabric/OFI, UCX, and RDMA concepts. * Experience with RoCEv2 and Ultra Ethernet. * Experience building cluster-scale performance test infrastructure. ## Description * Design and implement performance-critical features for CCL enablement on Cornelis Networks' fabrics. * Optimize distributed training performance across multi-node, multi-GPU configurations. * Improve GPU communication paths including GPU-direct transfers, IPC, and CPU/GPU synchronization. * Profile distributed AI workloads and identify bottlenecks across the software and hardware stack. * Tune AI frameworks such as PyTorch Distributed, TensorFlow/XLA, JAX, DeepSpeed, and Megatron-LM. * Develop benchmarks and microbenchmarks aligned with real model performance. * Contribute upstream to AI communication and distributed training projects. * Participate in design reviews, code reviews, CI, and long-term maintenance. * Prototype and validate Ultra Ethernet capabilities for AI collective communication. * Provide technical input for deployment considerations and performance validation. * Collaborate with kernel/driver, switch, performance, and systems teams. * Support advanced escalations by analyzing traces and providing robust fixes. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)