Principal Systems Engineer

NSCALE, LLC
Houston, TX, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$175,000.0 - $225,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Build Automation Computer Clusters Network Congestion Software Debugging Distributed Data Store Ethernet Firmware InfiniBand Network Layer PCI Express
+6 more
Performance Tuning Remote Direct Memory Access Ceph (Software) Graphics Processing Unit (GPU) High Performance Computing Nvme

Job description

We are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites.

You will own the full lifecycle of deployment-from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence.

This role combines deep hands-on expertise with system-level thinking and cross-functional leadership., Define the technical standards for node, rack, and full-cluster bringup. Lead large-scale GPU cluster deployments (multi-rack, multi-pod environments). Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads. Establish cluster-level acceptance criteria and validation frameworks.

Performance & Fabric Architecture

Tune and validate NCCL, RDMA, GPUDirect, and collective operations at scale. Identify and eliminate performance bottlenecks across hardware, topology, and firmware layers. Drive congestion control and fabric optimization strategies. Define performance benchmarking methodology for AI training workloads.

Deployment Strategy & Scalability

Design repeatable deployment models for multi-site expansion. Build automation frameworks for provisioning and cluster validation. Establish deployment SLAs, quality gates, and operational readiness standards. Reduce time-to-capacity while increasing reliability.

Technical Leadership

Serve as the escalation point for complex bringup and performance issues. Mentor senior engineers and shape infrastructure best practices. Influence hardware selection, rack topology, and data center design decisions. Partner with executive leadership on infrastructure scaling strategy.

Requirements

Do you have experience in System performance optimization?, 10+ years of experience in large-scale infrastructure or HPC environments. Proven experience bringing up large GPU clusters (hundreds+ GPUs). Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics). Strong understanding of server architecture (PCIe, NUMA, memory hierarchy). Experience debugging performance issues across compute and network layers. Strong automation and systems-level thinking.

Strongly Preferred

Experience scaling AI training clusters for frontier models. Experience with liquid cooling or ultra-high-density deployments. Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF). Experience defining infrastructure standards in a fast-growing organization.

About the company

We are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute.

As a startup, we move fast, operate with ownership, and expect technical leaders to define standards-not just follow them., Superclusters are brought online quickly, predictably, and at peak performance. Deployment processes scale from first cluster to multi-site expansion. Infrastructure becomes a competitive advantage.

  • You define the technical blueprint for how we scale AI infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · WWC 2025

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:20 min

Utilizing custom firmware for variable torque manipulation

Daniel Meilak Daniel Meilak +1 · WWC Europe 2026

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:12 min

Implementing automotive ethernet and connected remote vehicle telemetry applications

David Romić · WWC 2023

Videos

See all

Related articles

See all