Vice President of Infrastructure & Deployment (Remote or Hybrid)

Acasia Operations, Inc.
The Woodlands, United States of America
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 170K

Job location

The Woodlands, United States of America

Tech stack

Artificial Intelligence
Intelligent Platform Management Interface
BIOS
Cloud Computing
Computer Clusters
Nvidia CUDA
Data Centers
Dynamic Host Configuration Protocol
Linux
DNS
Ethernet
Firmware
InfiniBand
Storage Area Network (SAN)
Network Troubleshooting
Linux System Administration
Routing
Performance Tuning
Remote Direct Memory Access
Remote Infrastructure Management
Virtual Local Area Networks
AI Infrastructure
Graphics Processing Unit (GPU)
High Performance Computing
Computer Network Technologies
IT Architecture
Firewalls (Computer Science)
Kubernetes
Information Technology
Deployment Automation
Bare Metal
Slurm
Hardware Infrastructure
Docker

Job description

As Acasia expands across multiple data center markets, infrastructure design and technical execution become core differentiators. We are hiring a Vice President of Infrastructure & Deployment to serve as Acasia's senior infrastructure subject matter expert, technical authority, and escalation leader for GPU infrastructure design, deployment readiness, complex troubleshooting, ongoing support and production reliability., The Vice President of Infrastructure Engineering & Deployment is responsible for the successful delivery, implementation, commissioning, and operational readiness of Acasia's AI infrastructure across customer and data center environments.

This executive owns the complete infrastructure delivery lifecycle-from customer handoff after contract execution through deployment planning, installation, networking, cluster bring-up, validation, production acceptance, and ongoing infrastructure optimization. Their mission is to ensure every GPU cluster is delivered safely, on schedule, on budget, and performs to Acasia's standards before entering customer production.

The Vice President will build and lead Acasia's Infrastructure Delivery organization, managing field engineering teams responsible for implementing large-scale GPU clusters across multiple data centers and customer locations. They will establish deployment methodologies, technical standards, implementation playbooks, commissioning procedures, and quality controls while driving continuous improvements in deployment speed, consistency, and customer experience.

This leader will work hand-in-hand with customers, Sales, Customer Success, Engineering, Product, Operations, Procurement, OEM partners, networking vendors, data center operators, and technology partners to deliver world-class AI infrastructure at scale.

Beyond deployment execution, this executive serves as Acasia's highest-level infrastructure subject matter expert-guiding technical architecture, solving the company's most complex infrastructure challenges, mentoring engineering teams, and continuously improving Acasia's ability to deploy AI infrastructure faster and more reliably than anyone in the market.

Mission: Build the industry's fastest, most reliable GPU infrastructure deployment organization capable of taking customer environments from signed contract to production in weeks instead of months., Infrastructure Architecture & Technical Leadership

  • Own infrastructure architecture standards for GPU clusters across all Acasia deployments.
  • Define reference architectures for rack layouts, power distribution, cooling, networking, storage, monitoring, telemetry, and remote management.
  • Review customer infrastructure designs for scalability, resilience, serviceability, deployment efficiency, and lifecycle management.
  • Identify technical and deployment risks before implementation begins.
  • Establish repeatable infrastructure standards that simplify deployment while improving reliability and operational excellence.
  • Partner with Engineering, Product, Security, Sales, Customer Success, Procurement, OEMs, and customers to ensure infrastructure designs meet both technical and commercial objectives.

Networking & Cluster Integration

  • Own deployment and validation of high-performance AI networking environments.
  • Lead implementation of InfiniBand, Ethernet, RoCE, RDMA, NVLink, NVSwitch, GPUDirect, and NVIDIA networking technologies.
  • Oversee spine-leaf architecture deployment, east-west networking, switch configuration, storage networking, and cluster connectivity.
  • Establish standards for network validation, benchmarking, congestion analysis, latency optimization, and fabric performance tuning.
  • Validate cluster communication performance using NCCL, MPI, RDMA, and other AI infrastructure benchmarking tools.
  • Partner with NVIDIA and networking OEMs to optimize production environments.

Production Readiness & Commissioning

  • Define production acceptance standards for every customer deployment.
  • Own Factory Acceptance Testing (FAT), Site Acceptance Testing (SAT), cluster commissioning, and customer handoff.
  • Establish standards for hardware acceptance, BIOS and firmware consistency, GPU validation, driver installation, CUDA readiness, storage validation, telemetry, monitoring, and observability.
  • Lead burn-in procedures, stress testing, soak testing, redundancy validation, performance baselining, and workload validation before production acceptance.
  • Ensure every customer environment meets Acasia's technical standards before entering production.

Program Management & Deployment Execution

  • Build and manage deployment schedules for large-scale infrastructure implementations.
  • Coordinate execution across internal teams, customers, OEMs, contractors, and deployment partners.
  • Track milestones, critical path activities, dependencies, and implementation risks.
  • Lead executive deployment reviews and customer implementation meetings.
  • Develop standardized governance, reporting, and deployment management processes.
  • Drive continuous improvement in deployment efficiency, predictability, and execution quality.

Technical Escalation & Problem Resolution

  • Serve as Acasia's highest-level technical escalation point for complex infrastructure issues.
  • Lead cross-functional war-room responses during critical deployments or production-impacting incidents.
  • Troubleshoot issues spanning GPU hardware, Linux, networking, firmware, drivers, storage, environmental systems, and customer workloads.
  • Partner with Engineering, OEMs, vendors, data center providers, and customers on root cause analysis and permanent corrective actions.
  • Convert recurring issues into improved standards, tooling, automation, documentation, and deployment methodologies.

Leadership & Organizational Development

  • Build and lead Acasia's Infrastructure Delivery and Field Engineering organization.
  • Recruit, mentor, and develop world-class infrastructure engineers and deployment leaders.
  • Establish technical certification and training programs across the organization.
  • Build repeatable knowledge transfer mechanisms, technical playbooks, deployment standards, and troubleshooting frameworks.
  • Foster a culture of accountability, precision, technical excellence, customer obsession, and continuous improvement.

Vendor & Strategic Partner Management

  • Build executive technical relationships with NVIDIA, OEM partners, networking providers, storage vendors, and data center operators.
  • Evaluate new infrastructure technologies, deployment methodologies, and vendor capabilities.
  • Participate in technical due diligence for new data center markets and deployment strategies.
  • Hold strategic partners accountable during implementation and complex technical escalations.
  • Support commercial and procurement teams with technical guidance on infrastructure quality, lifecycle planning, supportability, and deployment risk., * Design and review infrastructure patterns for GPU servers, racks, power, cooling, networking, cabling, storage, monitoring, telemetry, and remote management.
  • Establish reference architectures for repeatable GPU cluster deployments.
  • Review proposed customer deployments for performance, reliability, maintainability, scalability, and supportability.
  • Identify technical risks in infrastructure designs before they become production issues.
  • Partner with engineering, product, security, customer success, sales, vendors, and data center providers to ensure infrastructure designs meet customer and business requirements.
  • Serve as Acasia's internal authority on GPU infrastructure architecture, data center deployment patterns, and production readiness.

Success Looks Like

Within the first 90 days, this person will:

  • Assess Acasia's current and planned GPU infrastructure designs.
  • Identify major architecture, deployment, supportability, and reliability risks.
  • Define initial reference architecture standards for GPU infrastructure deployments.
  • Establish technical validation criteria for customer-ready environments.
  • Create the first version of Acasia's infrastructure troubleshooting and escalation framework.
  • Coach local infrastructure teams on technical standards, diagnostic discipline, and escalation quality.
  • Improve visibility into the most important technical risks across deployed environments.

Within the first 6-12 months, this person will:

  • Establish Acasia's technical infrastructure standards across all data center markets.
  • Create repeatable infrastructure design patterns for GPU deployments.
  • Reduce complex incident resolution time through better architecture, tooling, troubleshooting frameworks, and team training.
  • Improve customer deployment quality and production readiness.
  • Build a stronger technical bench across local infrastructure teams.
  • Make Acasia more credible with enterprise customers, vendors, data center partners, and investors as a serious AI infrastructure operator.

Reporting Structure

This role is expected to report to the COO, CTO, or equivalent executive leader responsible for infrastructure execution, technical delivery, and customer reliability.

Pay: $140,000.00 - $170,000.00 per year

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, Information Technology, or a related technical discipline (Master's preferred).
  • 10+ years leading infrastructure engineering, field engineering, or large-scale infrastructure deployment organizations.
  • Demonstrated success deploying production AI or HPC infrastructure environments exceeding 500 GPUs; experience with deployments of 1,000+ GPUs preferred.
  • Extensive hands-on experience implementing GPU clusters, Linux environments, storage systems, and high-performance networking.
  • Deep expertise with rack-scale infrastructure, structured cabling, power distribution, cooling systems, and data center operations.
  • Strong Linux systems administration and troubleshooting expertise.
  • Strong networking expertise including InfiniBand, RoCE, Ethernet, switching, routing, VLANs, RDMA, DNS, DHCP, firewalls, and production network troubleshooting.
  • Experience leading Factory Acceptance Testing (FAT), Site Acceptance Testing (SAT), commissioning, burn-in, and production validation.
  • Experience serving as the senior technical escalation point for complex production infrastructure issues.
  • Strong project leadership, vendor management, and cross-functional execution skills.
  • Excellent communication and executive presentation skills.
  • Willingness to travel extensively to customer sites, deployment locations, OEM facilities, and data center markets., * NVIDIA Certified Professional (NCP) strongly preferred, ideally NVIDIA Certified Professional - AI Infrastructure (NCP-AII) or equivalent NVIDIA AI Infrastructure certification.
  • Experience deploying NVIDIA HGX, DGX, GB300, GB200, B300, B200, H200, H100, or equivalent AI infrastructure platforms.
  • Expert knowledge of NVIDIA Networking, NVLink, NVSwitch, CUDA, NCCL, GPUDirect Storage, BlueField DPUs, DCGM, NVML, Redfish, BMC/IPMI, iDRAC, and iLO.
  • Experience working directly with NVIDIA and leading OEM partners including Supermicro, Dell, HPE, Lenovo, ASUS, and GIGABYTE.
  • Experience with Kubernetes, Slurm, Docker, bare-metal provisioning, AI workload orchestration, and GPU resource management.
  • Experience in cloud infrastructure, hyperscale data centers, AI infrastructure companies, managed infrastructure providers, or HPC environments.
  • Professional certifications such as CCNP, CCIE, RHCE, RHCSA, or equivalent enterprise infrastructure credentials.
  • Own technical architecture standards for Acasia's GPU infrastructure environments across data center markets.

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • 401(k) matching
  • Vision insurance
  • Dental insurance, Base Salary: $250k-$300k

Target Bonus: 50%

Equity: 0.25-0.5%

Benefits: Health, Dental, Vision, 401k

Reporting: COO

Location & Travel: Remote or Hybrid: Regular travel to data center markets, customer deployment sites, vendor locations, and company operating sessions should be expected

About Acasia

Acasia builds, deploys, and operates high-performance GPU infrastructure for enterprise AI workloads. Our customers rely on Acasia to deliver production-grade GPU environments that are performant, reliable, scalable, and supportable in real-world data center conditions., * 401(k)

  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Vision insurance

Apply for this position