Data Center Lab Engineer

DDN, LLC
Colorado Springs, United States of America
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Colorado Springs, United States of America

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Intelligent Platform Management Interface
Bash
BIOS
System Configuration
Data Centers
Data Center Infrastructure Management (CIM)
Linux
RAID
File Systems
Distributed Data Store
Ethernet
Firmware
General Parallel File Systems
InfiniBand
Python
Laboratory Information Management Systems
Nagios
Ansible
Prometheus
Remote Infrastructure Management
TCP/IP
Virtual Local Area Networks
Zabbix
Motherboard
Scripting (Bash/Python/Go/Ruby)
Connectivity Problems
Grafana
Comptia Server+
Storage Technologies
Information Technology
Data Management
Machine Learning Operations
Network Server
Nvme

Job description

DDN is looking for an experienced Data Center Lab Engineer to own and operate an engineering lab environment hosting enterprise-class servers, storage platforms, high-speed networking, racks, power infrastructure, cabling, and related data center assets. The role will support product engineering, QA, performance benchmarking, customer issue reproduction, and infrastructure readiness for DDN storage and AI/HPC solutions.

The ideal candidate is highly hands-on, operationally disciplined, comfortable working in a fast-moving engineering lab, and able to coordinate across engineering, QA, IT, facilities, support, vendors, and leadership teams., * Manage day-to-day lab operations, including rack readiness, power availability, hardware placement, cabling hygiene, and access coordination.

  • Rack, stack, cable, label, move, decommission, and maintain servers, storage arrays, JBOD/JBOF shelves, switches, PDUs, UPS connections, and test equipment.
  • Maintain clean, safe, audit-ready lab infrastructure with accurate rack elevations, power maps, network maps, and cabling documentation.
  • Track rack space, power utilization, cooling constraints, and equipment placement to support current and future engineering needs.

Server, Storage & Hardware Support

  • Deploy and support enterprise servers, dense storage systems, disk shelves, NVMe platforms, GPU/AI systems, and high-performance testbeds.
  • Perform hardware diagnostics, component replacement, firmware/BIOS updates, burn-in checks, health validation, and basic system troubleshooting.
  • Support engineering and QA teams with hardware setup for development, validation, performance, HA/failover, scalability, and customer issue reproduction.
  • Handle HDD, SSD, NVMe, HBA, NIC, RAID controller, GPU, memory, CPU, PSU, fan, and motherboard-level replacement coordination as required.
  • Use BMC/IPMI/iDRAC/iLO/Redfish, vendor tools, system logs, and Linux utilities to validate platform health.

Networking, Cabling & Fabric Readiness

  • Install, trace, label, and troubleshoot Ethernet, InfiniBand, Fibre Channel, management, console, DAC, AOC, optical, QSFP, QSFP-DD, OSFP, SFP, and breakout cabling.
  • Understand port speed compatibility, optics/transceiver matching, cable types, breakout behavior, link negotiation, and switch-port mapping.
  • Support lab management networks, VLAN connectivity, high-speed data fabrics, and multi-node storage/compute cluster cabling.
  • Diagnose link failures, cable mismatches, port errors, optics issues, incorrect breakout usage, and physical-layer connectivity problems.

Power, Cooling, Safety & Rack Management

  • Manage rack-level power distribution including PDUs, redundant power feeds, C13/C14, C19/C20, IEC/NEMA connectors, and high-density equipment requirements.
  • Monitor rack power consumption, identify overload risks, and coordinate with facilities for power, cooling, UPS, and electrical safety needs.
  • Apply hot/cold aisle discipline, airflow management, blanking panels, cable routing, ESD practices, and safe handling of heavy infrastructure.
  • Support planned power-on, shutdown, maintenance, relocation, and decommissioning activities with minimal disruption.

Asset, Inventory & Vendor Lifecycle Management

  • Maintain accurate inventory for servers, storage, switches, disks, cables, optics, PDUs, spares, and lab accessories.
  • Track asset ownership, location, serial numbers, warranty, project allocation, lifecycle status, and decommissioning records.
  • Coordinate shipping, receiving, asset tagging, vendor engagement, RMA, spare management, and replacement workflows.
  • Improve lab documentation and operational processes to reduce dependency on tribal knowledge.

Requirements

  • 6-12+ years of hands-on experience in data center operations, engineering labs, server/storage hardware support, infrastructure operations, or similar roles.
  • Strong experience with rack/stack/cable operations, structured cabling, labeling, hardware installation, and physical lab management.
  • Working knowledge of enterprise servers from Dell, HPE, Supermicro, Lenovo, Intel/AMD platforms, or equivalent hardware ecosystems.
  • Experience with storage hardware including HDD, SSD, NVMe, JBOD, JBOF, RAID, storage shelves, dense storage enclosures, and high-throughput platforms.
  • Good understanding of Ethernet, InfiniBand, Fibre Channel basics, TCP/IP, VLANs, management networks, optics, transceivers, and high-speed cable standards.
  • Comfortable with Linux basics, remote management, console access, PXE/OS installation support, firmware updates, and hardware log collection.
  • Understanding of rack power, PDUs, UPS feeds, power redundancy, cooling, airflow, and electrical/data center safety practices.
  • Ability to maintain high-quality documentation, asset inventory, rack diagrams, cabling maps, and operational records.

Preferred / Good-to-Have Skills

  • Experience in HPC, AI/ML infrastructure, enterprise storage, product engineering labs, or large-scale test environments.
  • Exposure to DDN storage systems, Lustre, GPFS, BeeGFS, NVMe-oF, S3/Object storage, distributed storage, or parallel file systems.
  • Hands-on experience with NVIDIA/Mellanox InfiniBand, GPU servers, dense AI racks, NVLink/NVSwitch, high-speed Ethernet, or liquid-cooled platforms.
  • Basic automation/scripting using shell, Python, Ansible, or lab provisioning workflows.
  • Experience with monitoring or DCIM tools such as Grafana, Prometheus, Zabbix, Nagios, vendor management platforms, or asset management systems.
  • Relevant certifications such as CompTIA Server+, Network+, CCNA, RHCSA/Linux, or data center operations certifications., * Bachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or equivalent practical experience.
  • Proven ability to work independently in a physical lab/data center environment while coordinating with distributed engineering teams.
  • Willingness to work on-site in the lab and support planned maintenance windows or urgent infrastructure needs when required.

Behavioral Competencies

  • High ownership mindset with strong accountability for lab readiness and uptime.
  • Detail-oriented approach to cabling, labeling, inventory, documentation, and safety.
  • Strong troubleshooting mindset with urgency, persistence, and structured problem solving.
  • Effective communication with engineering, QA, IT, facilities, vendors, support, and leadership stakeholders.
  • Ability to prioritize competing requests in a fast-paced engineering environment., * Improved lab readiness, hardware availability, and turnaround time for engineering testbed setup.
  • Reduced downtime caused by hardware, cabling, power, cooling, or asset-tracking gaps.
  • Accurate and up-to-date inventory, rack elevations, cabling diagrams, power maps, and asset ownership records.
  • Fast response to infrastructure issues, vendor RMAs, component replacements, and customer reproduction setup needs.
  • A safer, cleaner, better-documented, and more scalable DDN engineering lab environment.

Apply for this position