> Markdown version of [/jobs/ext/2056067-gpu-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2056067-gpu-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU Systems Engineer - **Company:** CAPITAL TOWERS II, INC. - **Location:** New York, United States - **Experience:** Expert - **Salary:** $200,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Computer Clusters, Nvidia CUDA, Software Debugging, Linux, Network Topologies, Performance Tuning, Remote Direct Memory Access, Ansible, Graphics Processing Unit (GPU), Break Fix, Infrastructure Automation Frameworks, Puppet - **Published:** August 14, 2026 - **Apply:** https://www.dice.com/job-detail/2d1a20bb-a1ec-457a-8727-673bb61433b3 ## About the Role * 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments. * Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it. * Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance. * Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not. * Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on. * Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef. * Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer. * Clear communication. You will work daily with researchers, engineers, and vendors. Nice to Have: * Experience with the rest of the NVIDIA stack, such as NCCL and NVLink. ## Description Trading and research at the firm run around the clock and across the globe, and they run on infrastructure this team designs, builds, and operates. As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes. The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention. Responsibilities: * Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation. * Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them. * Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups. * Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing. * Own infrastructure projects end to end, from scope and design through implementation and long-term support. * Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)