Infrastructure Technical Specialist

Acasia Operations, Inc.
Phoenix, AZ, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$100,000.0 - $125,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Intelligent Platform Management Interface BIOS Cloud Computing Computer Clusters CompTIA Network+ Nvidia CUDA Computer Maintenance Computer Networks System Configuration Data Centers Dynamic Host Configuration Protocol
+24 more
Data Center Infrastructure Management (CIM) Linux Domain Name System (DNS) Network Interface Controllers Firmware Monitoring of Systems Issue Tracking Systems InfiniBand Networking Hardware IP Addressing Routing Remote Infrastructure Management Virtual Local Area Networks Virtualization Technology AI Infrastructure Graphics Processing Unit (GPU) Computer Network Operations Firewalls (Computer Science) Kubernetes Information Technology Slurm Hardware Infrastructure Network Server Docker

Job description

Acasia builds and operates GPU infrastructure for enterprise AI workloads. We help customers access high-performance compute by deploying, managing, and supporting GPU clusters in data center environments.

As demand for AI infrastructure grows, execution quality matters. Customers expect GPU environments that are installed correctly, configured cleanly, monitored actively, and maintained with urgency. This role sits on the front lines of that promise.

Role Summary

The Infrastructure Technical Specialist will help deploy, maintain, troubleshoot, and support Acasia’s GPU infrastructure inside data center environments. This is a hands-on technical operations role responsible for physical installation, hardware maintenance, system configuration support, basic network troubleshooting, incident response, documentation, and customer SLA support.

This person should be comfortable working in data centers, handling technical equipment, following strict procedures, and communicating clearly during high-pressure operational issues.

Key ResponsibilitiesGPU Infrastructure Deployment

  • Assist with the physical deployment of GPU servers, networking equipment, cabling, racks, PDUs, and related infrastructure.
  • Rack, stack, cable, label, and validate equipment according to Acasia standards.
  • Support equipment receiving, inventory tracking, staging, burn-in, and deployment readiness checks.
  • Coordinate with data center personnel, vendors, internal engineering teams, and logistics partners during installations.
  • Follow deployment runbooks and escalate deviations or blockers quickly.

Maintenance & Troubleshooting

  • Diagnose and troubleshoot server, GPU, power, cabling, storage, network, and connectivity issues.
  • Replace or coordinate replacement of failed components, including GPUs, NICs, drives, memory, power supplies, cables, fans, and other server parts.
  • Support firmware, BIOS, driver, and operating system configuration tasks under engineering guidance.
  • Assist with routine preventative maintenance and infrastructure health checks.
  • Document root causes, corrective actions, and recurring issues.

Customer SLA Support

  • Respond quickly to infrastructure incidents that may impact customer uptime, performance, or availability.
  • Support escalation workflows for customer-impacting issues.
  • Communicate status updates clearly to internal stakeholders.
  • Help maintain service reliability by following incident response procedures and documenting actions taken.
  • Participate in after-hours or on-call support as needed.

Configuration & Validation

  • Perform basic configuration and validation tasks for servers, GPUs, networking, storage, and monitoring tools.
  • Run diagnostic tests to confirm hardware and system readiness.
  • Validate that deployed infrastructure meets required standards before customer handoff.
  • Assist engineering teams with cluster bring-up, node validation, and performance checks.
  • Maintain accurate records of asset status, serial numbers, rack locations, configurations, and changes.

Process & Documentation

  • Follow standardized runbooks, checklists, and change control processes.
  • Create and update technical documentation, deployment notes, incident logs, and maintenance records.
  • Identify gaps in processes and recommend improvements.
  • Help build repeatable operational workflows as Acasia scales deployments across multiple data centers.

Requirements

Do you have experience in Safety protocol adherence?, * 2+ years of experience in data center operations, IT infrastructure, hardware support, network operations, systems administration, or technical field support.

  • Hands-on experience with servers, racks, cabling, switches, power distribution, and hardware troubleshooting.
  • Basic understanding of Linux systems.
  • Basic understanding of networking concepts such as IP addressing, VLANs, DNS, DHCP, routing, switching, and firewalls.
  • Strong attention to detail and ability to follow technical procedures.
  • Ability to work in data center environments, including lifting equipment, standing for extended periods, and following safety/security procedures.
  • Strong written and verbal communication skills.
  • Ability to work under pressure during incidents or customer-impacting outages.
  • Willingness to travel to data center sites as needed.

Preferred Qualifications

  • Experience with GPU servers, AI/HPC infrastructure, or high-density compute environments.
  • Experience with NVIDIA GPUs, CUDA environments, NVIDIA drivers, InfiniBand, RoCE, or high-performance networking.
  • Experience with Supermicro, Dell, HPE, Lenovo, ASUS, Gigabyte, or other enterprise server platforms.
  • Familiarity with data center tools such as DCIM systems, IPMI/BMC, iDRAC, iLO, Redfish, or remote management platforms.
  • Experience with monitoring tools, ticketing systems, and incident management workflows.
  • Familiarity with Kubernetes, Slurm, Docker, virtualization, or cloud infrastructure.
  • Certifications such as CompTIA Network+, Server+, Linux+, CCNA, or equivalent practical experience.

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • 401(k) matching
  • Vision insurance
  • Dental insurance, * Base: $100k-$125k
  • Bonus: 15%
  • Benefits: Health, Dental, Vision, 401k, * 401(k)
  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

Videos

See all

Related articles

See all