> Markdown version of [/jobs/ext/2709099-head-of-infrastructure](https://www.wearedevelopers.com/jobs/ext/2709099-head-of-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Head of Infrastructure - **Company:** MONTAUK CAPITAL, LLC - **Location:** New York, NY, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Intelligent Platform Management Interface, Ubuntu (Operating System), Computer Clusters, System Configuration, Data Centers, Linux, Memory Management, Networking Basics, Red Hat Enterprise Linux, Remote Infrastructure Management, Shell Script, AI Infrastructure, Delivery Pipeline, Break Fix, Bare Metal, Hardware Infrastructure - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/infrastructure-engineer-perimeter-compute-montauk-capital-8141355 ## About the Role You're a strong infrastructure engineer experienced with hardware deployment, data center environments, GPU selection, systems setup and design. You can manage the implementation details end to end and have ownership over the entire process. If AI infrastructure is your jam and you've built systems in production, we want to talk. * Strong infrastructure engineering experience and systems-level technical judgment * Experience deploying or managing compute infrastructure in real-world environments * Experience with data center, hardware, or GPU-based systems implementation * Experience owning GPU provisioning, hardware selection, and systems configuration * GPU scheduling and orchestration specifics: GPU type awareness, memory management, topology considerations, placement strategies for multi-GPU jobs, and fragmentation minimization * Bare-metal provisioning lifecycle: IPMI/Redfish, BMC-based remote management, PXE boot, and automated OS deployment workflows * On-board storage * Observability stack: distributed configuration and troubleshooting, plus monitoring, alerting, and tracing * Deployment planning, Hardware configuration, Operational troubleshooting * Linux systems depth: RHEL/Ubuntu, low-level troubleshooting, shell scripting * Security and operational best practices for bare metal * Deployment tooling at production scale * Networking fundamentals for inference workloads and OOB management * Startup / 0*1 DNA: You ship fast and communicate clearly. ## Description We're building the automation, orchestration, and monitoring layer that unifies disparate metro edge GPU nodes into a single software-managed compute platform. You'll own the definition, design, implementation, and execution of the hardware and infrastructure buildout, executing strategy across edge data center requirements, GPU selection, supply chain, technical implementation, operational maintenance and deployment as we scale. You'll take the foundational groundwork and execute across the entire hardware and infrastructure side of our company, transforming our roadmap into production scale compute for AI inferencing. You'll ensure the GPU clusters deliver on customer requirements, are highly-available, and will be the hands on expert for the hardware side of our business. Most importantly, you'll turn our high-level plans into real, technical execution, and will play a key role in making supply chain decisions about infrastructure and how we deploy, scale, and support it. What You'll Do * Own GPU infrastructure design and implementation details from planning through deployment * Own hardware selection, configuration, and deployment across early compute infrastructure * Help turn early technical groundwork into a functioning deployed system * Own the GPU roadmap we use to entice customers and build partnerships * Deploy, operate, and tune GPU clusters for both bare-metal and internal software stack * Own resilient networking implementation from each site to the cluster, including a robust OOB network for constant monitoring and management * Manage deployments at production scale * Interface with site ops on power, cooling, and connectivity * Build the automation and monitoring stack for distributed edge nodes * Own the supply chain for all infrastructure gear * Manage third party hardware vendors on provisioning, maintenance and break-fix support ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [Data Mining Accessibility](https://www.wearedevelopers.com/videos/802-data-mining-accessibility) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j)