Datacenter Infrastructure Specialist

Spectraforce
United States
24 days ago
Apply on leoforce.us
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Bash Shell Data Centers Linux InfiniBand Python (Programming Language) Network Troubleshooting Linux System Administration Uptime Remote Direct Memory Access Reliability Engineering
+14 more
Prometheus AI Infrastructure Datadog Graphics Processing Unit (GPU) High Performance Computing Computer Network Technologies Large Language Models Grafana Hardware Testing Infrastructure Automation Frameworks Bare Metal Hardware Infrastructure Docker Golang

Job description

Client is looking for a Datacenter Infrastructure Specialist to help manage and support its rapidly growing global fleet of high-density GPU infrastructure. This person will act as the technical bridge between hardware partners and internal engineering teams, ensuring that GPU servers, networking, Linux systems, and supporting infrastructure remain reliable and performant. The role is ideal for someone with a background in Datacenter Infrastructure, Systems Engineering, HPC, GPU infrastructure, or Site Reliability, particularly someone who has experience troubleshooting Linux, networking, NVIDIA GPUs, and high-performance compute environments. This is a hands-on infrastructure role focused on hardware validation, troubleshooting, uptime, incident management, automation, and partner support., * Validate new server and GPU hardware to ensure deployments meet client’s requirements for AI/ML workloads.

  • Monitor infrastructure health and identify performance degradation or potential failures before they impact customers.
  • Troubleshoot datacenter networking and infrastructure performance issues.
  • Support high-performance networking technologies such as RDMA, InfiniBand, and RoCE.
  • Install, configure, and troubleshoot the NVIDIA software stack, including GPU drivers and performance utilities.
  • Troubleshoot Linux systems at the OS, kernel, hardware, and performance layers.
  • Support multi-node GPU and HPC environments and help optimize system performance.
  • Assist with incident response and communicate technical issues clearly to internal teams, leadership, and infrastructure partners.
  • Provide technical guidance and support to client’s hardware and infrastructure partners.
  • Help develop automated operational workflows using AI/LLMs, scripts, and internal tools.
  • Create and maintain technical runbooks and troubleshooting procedures.
  • Help monitor and enforce infrastructure uptime and customer SLA requirements.

Requirements

  1. Infrastructure / Datacenter * 3-5 years of experience in: + Datacenter Engineering + Infrastructure Operations + Systems Engineering + Site Reliability / Infrastructure Reliability * Hands-on experience supporting physical or bare-metal infrastructure.

  2. Linux * Strong Linux system administration experience. * Comfortable troubleshooting: + OS issues + Kernel-level problems + Hardware issues + System performance + Drivers

  3. Datacenter Networking * Strong understanding of standard datacenter networking. * Experience with network performance troubleshooting. * RDMA, InfiniBand, or RoCE experience is highly preferred.

  4. GPU / AI Infrastructure * Hands-on experience with NVIDIA GPUs. * Experience installing/troubleshooting NVIDIA drivers and software utilities. * Understanding of multi-node GPU performance or distributed workloads.

  5. Containers * Experience with Docker/containerization.

  6. Communication * Strong written and verbal communication. * Ability to explain complex infrastructure, hardware, and networking issues to both technical and non-technical stakeholders.

Preferred Skills

  • HPC / High-Performance Computing experience.
  • Experience managing large-scale bare-metal HPC environments.
  • Startup or high-growth infrastructure experience.
  • Monitoring/observability tools:
  • Grafana
  • Prometheus
  • Datadog
  • Automation/scripting:
  • Python
  • Go/Golang
  • Bash
  • Experience working with LLMs or AI agents for infrastructure automation.
  • Experience building operational workflows and automation from scratch.
  • Experience working directly with hardware/infrastructure vendors or datacenter partners.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on leoforce.us
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

Videos

See all

Related articles

See all