Datacenter Infrastructure Specialist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
Client is looking for a Datacenter Infrastructure Specialist to help manage and support its rapidly growing global fleet of high-density GPU infrastructure. This person will act as the technical bridge between hardware partners and internal engineering teams, ensuring that GPU servers, networking, Linux systems, and supporting infrastructure remain reliable and performant. The role is ideal for someone with a background in Datacenter Infrastructure, Systems Engineering, HPC, GPU infrastructure, or Site Reliability, particularly someone who has experience troubleshooting Linux, networking, NVIDIA GPUs, and high-performance compute environments. This is a hands-on infrastructure role focused on hardware validation, troubleshooting, uptime, incident management, automation, and partner support., * Validate new server and GPU hardware to ensure deployments meet client’s requirements for AI/ML workloads.
- Monitor infrastructure health and identify performance degradation or potential failures before they impact customers.
- Troubleshoot datacenter networking and infrastructure performance issues.
- Support high-performance networking technologies such as RDMA, InfiniBand, and RoCE.
- Install, configure, and troubleshoot the NVIDIA software stack, including GPU drivers and performance utilities.
- Troubleshoot Linux systems at the OS, kernel, hardware, and performance layers.
- Support multi-node GPU and HPC environments and help optimize system performance.
- Assist with incident response and communicate technical issues clearly to internal teams, leadership, and infrastructure partners.
- Provide technical guidance and support to client’s hardware and infrastructure partners.
- Help develop automated operational workflows using AI/LLMs, scripts, and internal tools.
- Create and maintain technical runbooks and troubleshooting procedures.
- Help monitor and enforce infrastructure uptime and customer SLA requirements.
Requirements
-
Infrastructure / Datacenter * 3-5 years of experience in: + Datacenter Engineering + Infrastructure Operations + Systems Engineering + Site Reliability / Infrastructure Reliability * Hands-on experience supporting physical or bare-metal infrastructure.
-
Linux * Strong Linux system administration experience. * Comfortable troubleshooting: + OS issues + Kernel-level problems + Hardware issues + System performance + Drivers
-
Datacenter Networking * Strong understanding of standard datacenter networking. * Experience with network performance troubleshooting. * RDMA, InfiniBand, or RoCE experience is highly preferred.
-
GPU / AI Infrastructure * Hands-on experience with NVIDIA GPUs. * Experience installing/troubleshooting NVIDIA drivers and software utilities. * Understanding of multi-node GPU performance or distributed workloads.
-
Containers * Experience with Docker/containerization.
-
Communication * Strong written and verbal communication. * Ability to explain complex infrastructure, hardware, and networking issues to both technical and non-technical stakeholders.
Preferred Skills
- HPC / High-Performance Computing experience.
- Experience managing large-scale bare-metal HPC environments.
- Startup or high-growth infrastructure experience.
- Monitoring/observability tools:
- Grafana
- Prometheus
- Datadog
- Automation/scripting:
- Python
- Go/Golang
- Bash
- Experience working with LLMs or AI agents for infrastructure automation.
- Experience building operational workflows and automation from scratch.
- Experience working directly with hardware/infrastructure vendors or datacenter partners.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
7 Cloud Computing Trends Coming in 2025 for Developers
Highest Paying Tech Companies for Developers
Top Big Data Technologies That You Need to Know