> Markdown version of [/jobs/ext/740051-data-center-operations-engineer](https://www.wearedevelopers.com/jobs/ext/740051-data-center-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Center Operations Engineer - **Company:** Nmk Global Inc. - **Location:** California City, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Address Resolution Protocols, Artificial Intelligence, Big Data, System Configuration, Data Centers, Linux, Microprocessors, RAID, Network Interface Controllers, Firmware, Internet Control Message Protocol, Issue Tracking Systems, InfiniBand, Internet Protocol, Subnetting, OSI Models, Linux System Administration, Linux Commands, Linux Servers, Simple Mail Transfer Protocols, Nagios, Networking Basics, TCP/IP, Transmission Control Protocol (TCP), Trivial File Transfer Protocols, Enterprise Data Management, Network Routers, Scripting, Graphics Processing Unit (GPU), File Transfer Protocol (FTP), Performance Testing, System Availability, Hardware Infrastructure, Terminal Servers - **Published:** June 29, 2026 - **Apply:** https://www.dice.com/job-detail/5d1e1d3d-3f99-4a5d-8e9e-371ac91a34a2 ## About the Role * 5+ years of experience in Data Center Operations or Infrastructure Engineering. * Strong hands-on experience with Linux system administration, troubleshooting, and performance validation. * Experience with Linux command-line utilities and Bash/Shell scripting. * Hands-on experience deploying and configuring GPU servers in clustered environments. * Experience with GPU cluster bring-up, driver installation, and system-level configuration. * Strong knowledge of InfiniBand networking, including switch configuration, subnet management, and troubleshooting. * Experience performing end-to-end GPU testing in InfiniBand-based clusters. * Solid understanding of networking fundamentals, including TCP/IP, OSI Model, ARP, ICMP, TCP, UDP, SMTP, FTP, and TFTP. * Experience installing, configuring, and troubleshooting routers, switches, and terminal servers. * Hands-on experience with server hardware installation, rack and stack, cabling, CPUs, memory, HDDs, RAID controllers, NICs, and firmware upgrades. * Experience with fiber and copper cabling, IP networking, and SAN infrastructure. * Experience supporting data center deployments, migrations, hardware refreshes, and expansion projects. * Experience using monitoring and alerting tools to identify and resolve infrastructure issues. * Experience working with ticketing systems while meeting SLA requirements. * Strong documentation skills for operational procedures, system configurations, and technical runbooks. * Excellent troubleshooting, communication, and organizational skills. * Ability to work in a fast-paced production environment and participate in on-call rotations. Preferred Skills: * Experience supporting HPC, AI, or large-scale GPU environments. * Experience with NVIDIA GPU platforms and Mellanox/InfiniBand technologies. * Experience with data center monitoring solutions. * Experience supporting large-scale data center build-outs and infrastructure refresh programs. * Familiarity with automation or scripting for operational tasks. ## Description We are seeking a Data Center Operations Engineer with strong hands-on experience supporting enterprise data center infrastructure, Linux systems, GPU server deployments, and InfiniBand networking. The ideal candidate will have expertise in installing, configuring, troubleshooting, and maintaining data center hardware and infrastructure while supporting HPC/AI environments and ensuring high availability of critical systems., This role requires excellent troubleshooting skills, experience with GPU cluster deployments, InfiniBand fabrics, Linux administration, networking, and data center operations. The engineer will work closely with infrastructure, operations, and engineering teams to support deployments, maintenance activities, and continuous operational improvements., * Provide operational support for data center deployments, maintenance, and repair activities. * Install, configure, test, and maintain Linux servers and GPU infrastructure. * Deploy, configure, and validate GPU servers and clustered environments. * Perform InfiniBand fabric bring-up, switch configuration, subnet management, and troubleshooting. * Install and maintain server hardware, including CPUs, memory, storage, RAID components, and network adapters. * Configure and troubleshoot routers, switches, terminal servers, and out-of-band management devices. * Perform daily health checks of Linux systems, networking, and infrastructure components. * Support data center build-outs, hardware refreshes, migrations, and expansion projects. * Coordinate with vendors for hardware installation, diagnostics, replacement, and warranty support. * Monitor infrastructure using monitoring and alerting tools, ensuring timely incident resolution. * Maintain operational documentation, technical procedures, and runbooks. * Participate in incident response, maintenance windows, and on-call support rotations. * Collaborate with cross-functional global teams to ensure reliable, secure, and scalable infrastructure operations. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)