> Markdown version of [/jobs/ext/643957-systems-engineer-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/643957-systems-engineer-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Engineer - AI Infrastructure - **Company:** CLOCKWORK SYSTEMS, INC. - **Location:** United States - **Experience:** Expert - **Salary:** $150,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, C++ (Programming Language), Computer Clusters, Profiling, Nvidia CUDA, Databases, Software Debugging, Device Drivers, Distributed Data Store, Distributed Systems, Fault Tolerance, InfiniBand, Remote Direct Memory Access, Software Engineering, Pytorch, Gpu Programming - **Published:** June 25, 2026 - **Apply:** https://www.dice.com/job-detail/83462549-0721-4b37-ae7f-ab4c28cd291f ## About the Role Required: Systems building experience, * Kernel subsystems, device drivers, or OS-level components * Distributed storage, databases, or coordination systems * Runtimes, profilers, or performance tooling * Network stacks, protocols, or high-performance I/O systems * Large-scale infrastructure at the systems layer Core technical skills: * Strong C/C++ in systems contexts (not just application code) * Deep understanding of concurrency, memory models, and failure modes * Experience reasoning about distributed system behavior: consistency, ordering, partial failures * Comfortable reading and modifying large, unfamiliar codebases Nice to have: * GPU programming (CUDA) or GPU systems experience * High-performance networking (RDMA, InfiniBand) * ML framework or runtime internals * Cluster scheduling or orchestration systems We believe strong systems engineers pick up domain-specific tools quickly. We value your ability to reason about complex systems over checkbox familiarity with our specific stack. Senior Expectations * Lead design of significant system components * Navigate ambiguity and define technical direction * Mentor engineers and raise team capabilities * 8+ years building systems software ## Description We're building infrastructure for fault-tolerant, high-performance distributed GPU training. You'll work at the intersection of GPU systems, high-speed networking, and distributed coordination-designing and implementing systems that run at scale. This is a systems building role. You'll dig into internals, understand why things break under pressure, and design solutions that handle the messy reality of distributed systems. What You'll Do * Design and implement low-level systems software for GPU clusters * Work with internals of frameworks like PyTorch, NCCL, CUDA runtime-not as a user, but modifying and extending them * Build components that make large-scale GPU training more reliable and efficient * Debug complex distributed/concurrent systems where failures are subtle and non-deterministic * Own systems end-to-end: from design through production ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)