> Markdown version of [/jobs/ext/2058632-datacenter-infrastructure-specialist](https://www.wearedevelopers.com/jobs/ext/2058632-datacenter-infrastructure-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Datacenter Infrastructure Specialist - **Company:** Spectraforce - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Bash Shell, Data Centers, Linux, InfiniBand, Python (Programming Language), Network Troubleshooting, Linux System Administration, Uptime, Remote Direct Memory Access, Reliability Engineering, Prometheus, AI Infrastructure, Datadog, Graphics Processing Unit (GPU), High Performance Computing, Computer Network Technologies, Large Language Models, Grafana, Hardware Testing, Infrastructure Automation Frameworks, Bare Metal, Hardware Infrastructure, Docker, Golang - **Published:** August 14, 2026 - **Apply:** http://leoforce.us/Careers/Spectraforce/JobDetails.html?jobid=a063d232-a82c-4f47-bed4-0da374171947&OrgId=1&UserId=5378 ## About the Role 1. Infrastructure / Datacenter * 3-5 years of experience in: + Datacenter Engineering + Infrastructure Operations + Systems Engineering + Site Reliability / Infrastructure Reliability * Hands-on experience supporting physical or bare-metal infrastructure. 2. Linux * Strong Linux system administration experience. * Comfortable troubleshooting: + OS issues + Kernel-level problems + Hardware issues + System performance + Drivers 3. Datacenter Networking * Strong understanding of standard datacenter networking. * Experience with network performance troubleshooting. * RDMA, InfiniBand, or RoCE experience is highly preferred. 4. GPU / AI Infrastructure * Hands-on experience with NVIDIA GPUs. * Experience installing/troubleshooting NVIDIA drivers and software utilities. * Understanding of multi-node GPU performance or distributed workloads. 5. Containers * Experience with Docker/containerization. 6. Communication * Strong written and verbal communication. * Ability to explain complex infrastructure, hardware, and networking issues to both technical and non-technical stakeholders. Preferred Skills * HPC / High-Performance Computing experience. * Experience managing large-scale bare-metal HPC environments. * Startup or high-growth infrastructure experience. * Monitoring/observability tools: + Grafana + Prometheus + Datadog * Automation/scripting: + Python + Go/Golang + Bash * Experience working with LLMs or AI agents for infrastructure automation. * Experience building operational workflows and automation from scratch. * Experience working directly with hardware/infrastructure vendors or datacenter partners. ## Description Client is looking for a Datacenter Infrastructure Specialist to help manage and support its rapidly growing global fleet of high-density GPU infrastructure. This person will act as the technical bridge between hardware partners and internal engineering teams, ensuring that GPU servers, networking, Linux systems, and supporting infrastructure remain reliable and performant. The role is ideal for someone with a background in Datacenter Infrastructure, Systems Engineering, HPC, GPU infrastructure, or Site Reliability, particularly someone who has experience troubleshooting Linux, networking, NVIDIA GPUs, and high-performance compute environments. This is a hands-on infrastructure role focused on hardware validation, troubleshooting, uptime, incident management, automation, and partner support., * Validate new server and GPU hardware to ensure deployments meet client's requirements for AI/ML workloads. * Monitor infrastructure health and identify performance degradation or potential failures before they impact customers. * Troubleshoot datacenter networking and infrastructure performance issues. * Support high-performance networking technologies such as RDMA, InfiniBand, and RoCE. * Install, configure, and troubleshoot the NVIDIA software stack, including GPU drivers and performance utilities. * Troubleshoot Linux systems at the OS, kernel, hardware, and performance layers. * Support multi-node GPU and HPC environments and help optimize system performance. * Assist with incident response and communicate technical issues clearly to internal teams, leadership, and infrastructure partners. * Provide technical guidance and support to client's hardware and infrastructure partners. * Help develop automated operational workflows using AI/LLMs, scripts, and internal tools. * Create and maintain technical runbooks and troubleshooting procedures. * Help monitor and enforce infrastructure uptime and customer SLA requirements. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023)