Senior Manager, AI Clusture Deployment

5C DATA CENTERS USA INC.
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$145,000.0 - $175,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence JIRA Cloud Computing Computer Clusters Data Centers Data Center Infrastructure Management (CIM) Linux Ethernet Hardware Design Monitoring of Systems InfiniBand Linux System Administration
+17 more
Machine Learning Network Architecture Performance Tuning Remote Direct Memory Access Software Deployment AI Infrastructure Data Storage Technologies Performance Testing IT Architecture AI Platforms Kubernetes Infrastructure Automation Frameworks Storage Technologies Information Technology Deployment Automation Hardware Infrastructure Data Delivery

Job description

We are seeking an experienced Senior Manager of AI Cluster Deployment to lead the planning, deployment, integration, and operational readiness of large-scale AI infrastructure environments. This role is responsible for delivering production-grade GPU clusters that support AI training, inference, and high-performance computing workloads across cloud, hybrid, and on-premises environments.

The ideal candidate brings deep technical expertise in GPU infrastructure, networking, storage, automation, and datacenter deployment, combined with strong program leadership and cross-functional execution skills. This leader will oversee end-to-end AI cluster deployment initiatives, including hardware integration, rack-and-stack operations, provisioning automation, performance validation, and operational handoff.

The role requires hands-on familiarity with modern AI infrastructure tooling and architectures, including Canonical MaaS, VAST Data storage platforms, and both InfiniBand and Ethernet-based GPU networking fabrics., AI and GPU Cluster DeploymentDelivery

Oversee and partake in deployment and integration of GPU-based compute platforms from NVIDIA and other accelerator vendors

Lead and participate in end-to-end logical deployment of large-scale AI and GPU clusters in state of the art datacenters.

Manage deployment programs spanning compute, storage, networking, power, cooling, and automation layers.

Participate in cluster architecture review for AI training, inference and distributed compute workloads

Coordinate rack-and-stack and cabling sequencing, network deployment, burn-in testing, and cluster validation.Validate deployment readiness, topology consistency, GPU fabric performance, acceptance testing, and operational turnover processes.

Establish repeatable and documented deployment methodologies and scalable operational standards.

NetworkingFabric Management

Lead deployment and operational validation of high-performance GPU interconnects using InfiniBand and Ethernet GPU fabric architectures

Ensure proper implementation of: spile-leaf architectures, RDMA, network telemetry and performance tuning

Coordinate closely with network engineering teams on topology implementation and performance optimization.

StorageData Infrastructure

Coordinate with storage engineering teams on deployment and integration of high-performance storage environments supporting AI workloads.

Ensure successful implementation and operational optimization of data storage platforms

Validate storage throughput, latency, and GPU data delivery performance.

AutomationProvisioning

Lead infrastructure automation initiatives for cluster provisioning and lifecycle management.

Manage deployment tooling and orchestration platforms including:

o Infrastructure-as-Code frameworks

o Automated imaging and provisioning systems (e.g. Canonical MaaS)

o Cluster monitoring and observability tools

Drive standardization and deployment automation to improve speed, reliability, and repeatability.

LeadershipProgram Management

Build and lead high-performing technical deployment and infrastructure engineering teams.

Partner with datacenter operations, hardware vendors, networking teams, and AI platform engineering groups.

Establish strong Project Management Office (PMO) partnership while driving consistent, accurate project updates across the team and systems (e.g. Jira)

Develop operational procedures, documentation, and deployment best practices.

Mentor engineers and technical leads across infrastructure domains.

Requirements

Do you have experience in System deployment?, Bachelor’s degree in Computer Science, Engineering, Information Technology, or related field (or equivalent experience).

10+ years of infrastructure engineering or datacenter deployment experience.

5+ years leading deployment or operations teams supporting large-scale AI, HPC, or GPU infrastructure.

Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments.

Strong expertise with:

o Canonical MaaS

o Data storage platforms

o InfiniBand and Ethernet GPU fabrics

o Network architecture

o Linux systems administration

o GPU server architectures

Strong understanding of:

o RDMA and RoCE networking

o High-performance storage architectures

o Cluster automation and provisioning

o Datacenter infrastructure operations

Proven ability to manage complex cross-functional infrastructure deployment programs.

Preferred Qualifications

Experience deploying NVIDIA DGX SuperPOD or similar AI infrastructure solutions.

Familiarity with:

o NVIDIA networking technologies

o Spectrum-X or Quantum platforms

o AI model training infrastructure

o Liquid cooling environments

o DCIM and observability platforms

Experience in hyperscale, cloud, or AI infrastructure environments.

Certifications in networking, Linux, Kubernetes, or cloud infrastructure are a plus.

Key Competencies

Technical leadership

Infrastructure architecture

Program execution

Cross-functional collaboration

Vendor and stakeholder management

Problem-solving under operational pressure

Process improvement and automation

Excellent communication and documentation skills

Benefits & conditions

5C Data Centers USA, Inc. United States Remote $145,000 - $175,000 a year - Full-time

About the company

5C Data Centers is an equal opportunity employer. We evaluate all qualified applicants without regard to race, religion, gender, age, national origin, disability, sexual orientation, veteran status, or other protected status.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Core infrastructure components required for an AI factory

Thomas Schmidt Thomas Schmidt · WWC 2024

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

39 sec

Deploying an immediate AI center of excellence via superpods

Thomas Schmidt Thomas Schmidt · WWC 2024

2:12 min

Implementing automotive ethernet and connected remote vehicle telemetry applications

David Romić · WWC 2023

Videos

See all

Related articles

See all