> Markdown version of [/jobs/ext/607669-principal-systems-engineer](https://www.wearedevelopers.com/jobs/ext/607669-principal-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Systems Engineer - **Company:** NSCALE, LLC - **Location:** Houston, TX, United States - **Experience:** Expert - **Salary:** $175,000.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Build Automation, Computer Clusters, Network Congestion, Software Debugging, Distributed Data Store, Ethernet, Firmware, InfiniBand, Network Layer, PCI Express, Performance Tuning, Remote Direct Memory Access, Ceph (Software), Graphics Processing Unit (GPU), High Performance Computing, Nvme - **Published:** June 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dc24e3917b27c932 ## About the Role Do you have experience in System performance optimization?, 10+ years of experience in large-scale infrastructure or HPC environments. Proven experience bringing up large GPU clusters (hundreds+ GPUs). Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics). Strong understanding of server architecture (PCIe, NUMA, memory hierarchy). Experience debugging performance issues across compute and network layers. Strong automation and systems-level thinking. Strongly Preferred Experience scaling AI training clusters for frontier models. Experience with liquid cooling or ultra-high-density deployments. Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF). Experience defining infrastructure standards in a fast-growing organization. ## Description We are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites. You will own the full lifecycle of deployment-from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence. This role combines deep hands-on expertise with system-level thinking and cross-functional leadership., Define the technical standards for node, rack, and full-cluster bringup. Lead large-scale GPU cluster deployments (multi-rack, multi-pod environments). Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads. Establish cluster-level acceptance criteria and validation frameworks. Performance & Fabric Architecture Tune and validate NCCL, RDMA, GPUDirect, and collective operations at scale. Identify and eliminate performance bottlenecks across hardware, topology, and firmware layers. Drive congestion control and fabric optimization strategies. Define performance benchmarking methodology for AI training workloads. Deployment Strategy & Scalability Design repeatable deployment models for multi-site expansion. Build automation frameworks for provisioning and cluster validation. Establish deployment SLAs, quality gates, and operational readiness standards. Reduce time-to-capacity while increasing reliability. Technical Leadership Serve as the escalation point for complex bringup and performance issues. Mentor senior engineers and shape infrastructure best practices. Influence hardware selection, rack topology, and data center design decisions. Partner with executive leadership on infrastructure scaling strategy. ## Related Videos - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Building a hypercar from scratch](https://www.wearedevelopers.com/videos/607-building-a-hypercar-from-scratch) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)