Senior AI GPU Deployment Engineer

5C DATA CENTERS USA INC.
Springfield, OH, United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$120,000.0 - $150,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computing Platforms Bash Shell BIOS Computer Clusters Configuration Management Data Centers Ethernet Network Interface Controllers Firmware InfiniBand Subnetting
+11 more
Python (Programming Language) Linux System Administration Remote Direct Memory Access Ansible SQL Databases AI Infrastructure Graphics Processing Unit (GPU) High Performance Computing Infrastructure Automation Frameworks Information Technology Hardware Infrastructure

Job description

We are seeking an experienced Senior AI GPU Deployment Engineer to plan, deploy, and operationalize large-scale GPU AI infrastructure environments. This role delivers production-grade GPU clusters supporting AI training, inference, and high-performance computing workloads in our hyperscale data centers., We’re looking for a Senior AI GPU Deployment Engineer to join our Cloud Ops team in the United States. This person will share our company values and play an important role in supporting our continued growth.

What You Will Do

  • Deploy integrate and validate multi-rack GPU-based compute platform deployments
  • Deploy fabric configuration engines (Subnet Manager), observability platforms (UFM) and validate interconnect and fabric performance (nccl)
  • Collaborate with network engineering team on topology implementation and optimization and storage engineering team on deployment and integration of high-performance storage environments supporting AI workloads (e.g. VAST Data)
  • Configure settings and manage firmware updates for GPUs, NICs, BMC, BIOS and other components across large-scale clusters
  • Contribute to infrastructure-as-code automation development for cluster provisioning and lifecycle management
  • Contribute to improving and documenting repeatable deployment methodologies and scalable operational standards
  • Query and analyze deployment outcomes using SQL for diagnostics and operational reporting

Requirements

The ideal candidate brings deep technical expertise in GPU infrastructure, network fabrics, storage, automation, and Linux systems administration, combined with strong execution and troubleshooting skills., * Bachelor’s degree in Computer Science, Engineering, IT, or related field (or equivalent experience)

  • 5+ years of infrastructure engineering or datacenter deployment experience
  • 3+ years deploying large-scale AI, HPC, or GPU infrastructure
  • Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments
  • Strong expertise with:
  • GPU architectures
  • InfiniBand (NDR/XDR) and Ethernet GPU fabrics (Spectrum-X)
  • NVLink, NVSwitch, and GPU-direct technologies
  • Canonical MaaS and automated provisioning systems
  • VAST Data or similar high-performance storage platforms
  • Linux systems administration for HPC/AI workloads
  • Infrastructure-as-Code and configuration management (Ansible)
  • Python, Shell, and SQL for infrastructure automation and diagnostics
  • Strong understanding of:
  • RDMA, RoCE, and lossless Ethernet fabrics
  • Cluster automation, observability, and lifecycle management

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Pension plan

About the company

Why Join 5C Data Centers?

At 5C, we believe great people build great companies. You’ll build a rewarding career while helping shape the future of digital infrastructure - one of the fastest-growing industries in the world.

Career Growth

Build a rewarding career in one of the world’s fastest-growing industries.

Industry Leadership

Help power the infrastructure behind AI and high-performance computing.

Entrepreneurial Culture

Your ideas matter - we empower employees to help shape our future.

Comprehensive Rewards

Competitive pay plus meaningful, lasting impact on the work you do.

Life at 5C

We’re more than a workplace - we’re a team of builders, innovators, and problem-solvers united by a shared purpose: creating infrastructure that powers the technologies transforming our world. Your voice matters here, and we encourage fresh ideas at every level.

Ready to Apply?

If this opportunity sounds like the right fit for you, we’d love to hear your story. Apply today and discover what your future could look like at 5C Data Centers.

5C Data Centers is an equal opportunity employer.

We celebrate diversity and are committed to creating an inclusive environment where everyone can thrive. 5C evaluates qualified applicants without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity or expression, disability status, or any other legally protected characteristic.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · WWC 2024

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · WWC 2025

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · WWC Europe 2026

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · WWC 2025

Videos

See all

Related articles

See all