> Markdown version of [/jobs/ext/2608869-data-center-field-engineer](https://www.wearedevelopers.com/jobs/ext/2608869-data-center-field-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Center Field Engineer - **Company:** Aquila Hash, Inc. - **Location:** Barker, NY, United States - **Experience:** Starter - **Salary:** $65,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Intelligent Platform Management Interface, Computer Clusters, Data Centers, Linux, General-Purpose Computing on Graphics Processing Units, Issue Tracking Systems, InfiniBand, Networking Hardware, Network Architecture, Remote Direct Memory Access, Broadcom, Prometheus, TCP/IP, Zabbix, Network Switches, Computer Network Technologies, Grafana, Break Fix, Data Center Networking, Information Technology, Hardware Infrastructure, Network Server, Server Operating Systems & Platforms - **Published:** August 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=74067ae9700ad5ee ## About the Role * Degree or relevant educational background in Computer Science, Information Technology, Engineering, Telecommunications or a related field. * 1-2 years of hands-on experience in server, network, data center or infrastructure operations and maintenance. * Ability to lead and coordinate a team of 3-4 technicians during an assigned shift, prioritize operational tasks and serve as the primary on-shift escalation point. * Hands-on knowledge of x86 server hardware and Linux operating systems. * Working knowledge of TCP/IP and data center networking. * Hands-on ability to diagnose server hardware failures and replace field-replaceable components. * Familiarity with server BMC/OOB management interfaces such as iDRAC, IPMI or equivalent. * Ability to interpret monitoring alerts, hardware logs and system information to identify and troubleshoot infrastructure problems. * Experience with monitoring platforms such as Zabbix, Prometheus and Grafana is preferred. * Experience with GPU servers, RDMA, InfiniBand, RoCE or other high-performance networking technologies is preferred. * Good written and verbal communication skills, including the ability to prepare incident reports, RCA documentation, shift reports and technical emails. * Ability to communicate effectively with OEM support teams through email, ticketing systems and phone calls. * Strong sense of ownership, good organizational skills and the ability to make sound operational decisions during incidents. * Willing and able to work 24×7 rotating shifts in an on-site data center environment. Preferred Qualifications * Experience operating large-scale AI GPU clusters. * Experience supporting NVIDIA or AMD GPU server platforms. * Troubleshooting experience with NVIDIA/Mellanox, Broadcom or similar high-performance networking equipment. * Experience with Dell, Supermicro or other enterprise GPU/server platforms and OEM support processes. * RHCE, CCNA/CCNP or similar industry certifications. * Experience with ITIL-based incident, problem and change-management processes. * Previous experience coordinating technicians, shift activities or data center operations is a plus., * How many years of hands-on experience do you have working in data center operations, server hardware, or network infrastructure? * Have you previously led or coordinated data center technicians during a shift? If yes, how many people? * Are you willing and able to work on-site in a 24×7 rotating shift schedule, including nights and weekends? ## Description The AI GPU Cluster Shift Engineer is responsible for leading on-site operations and maintenance activities for large-scale AI GPU computing clusters during an assigned shift. Each Shift Engineer will lead and coordinate a team of approximately 3-4 Data Center / GPU Operations Technicians, ensuring effective 24×7 operational coverage of GPU servers, network infrastructure and related data center equipment. The Shift Engineer serves as the technical lead and primary operational point of contact during the shift, responsible for monitoring infrastructure health, coordinating incident response, assigning work to technicians, troubleshooting escalated issues, managing vendor support cases and ensuring a complete handover to the next shift., * Lead and coordinate 3-4 technicians during the assigned shift and manage day-to-day operational activities and work assignments. * Monitor the health and status of large-scale AI GPU clusters, including GPU/CPU servers, network switches, optical transceivers and associated infrastructure. * Review monitoring alerts, identify infrastructure abnormalities and determine appropriate troubleshooting and escalation actions. * Provide technical guidance to technicians for hardware troubleshooting, component replacement, rack-and-stack, cabling, server recovery and other on-site activities. * Perform hands-on troubleshooting of server, GPU, network and hardware issues that cannot be resolved by the technician team. * Coordinate incident response during the shift and ensure incidents are properly documented, tracked, escalated and communicated to relevant teams. * Serve as the primary on-shift escalation point and coordinate with engineering teams, management and other stakeholders as required. * Coordinate with OEM and hardware vendors for technical support, hardware replacement and RMA activities, and track cases through resolution. * Manage operational changes according to established change-management procedures and Standard Operating Procedures (SOPs). * Ensure technicians follow approved SOPs, safety requirements, change procedures and data center operating policies. * Verify completion and quality of maintenance activities, hardware replacements, cabling changes and other work performed during the shift. * Maintain accurate shift logs, incident records, maintenance records and operational documentation. * Conduct a structured shift handover to the incoming Shift Engineer, including active incidents, degraded equipment, pending maintenance, open vendor/RMA cases and other outstanding operational issues. * Participate in incident reviews and Root Cause Analysis (RCA), and recommend improvements to SOPs and operational processes. * Support spare-parts management and ensure critical replacement components are properly tracked and available for operations. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers)