> Markdown version of [/jobs/ext/137176-head-of-infrastructure-support](https://www.wearedevelopers.com/jobs/ext/137176-head-of-infrastructure-support). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Head of Infrastructure Support - **Company:** NSCALE, LLC - **Location:** Madison, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Bash Shell, Cloud Computing, Continuous Integration, Data Centers, Software Debugging, Linux, Distributed Systems, InfiniBand, Python (Programming Language), Network Layer, Routing, OpenStack, Queue Management Systems, Remote Direct Memory Access, Site Reliability Engineering Practices, Ansible, Runbook, Virtual Local Area Networks, Load Balancing, Kubernetes, Infrastructure Automation Frameworks, Terraform - **Published:** May 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8b7a38bb63657f5d ## About the Role Do you have experience in Stakeholder relationship building?, * You're comfortable problem solving & making decisions on complex topics with high levels of ambiguity in a results-driven environment. * You're comfortable influencing without authority and exceptional at building relationships with senior stakeholders across the business to get things done. * You have the understanding and skillset to grasp technical concepts and problems quickly. * You have strong analytical skills. * You're a doer who is extremely organised and diligent. * You're a self starter, curious, and quick to learn, knowing what questions to ask to get up to speed quickly., * Adaptable to customer-driven demands, including out-of-hours support and travel for onsite technical work * Disciplined, organised, and self-motivated, with the ability to lead, mentor, and support engineers in a fast-paced environment * Strong leadership mindset with a bias for decisive action, accountability, and continuous improvement * Experience leading or managing engineers in an operational support environment, including performance, development, and day-to-day team oversight * Experience owning team workload, prioritisation, and service delivery to meet SLAs * Excellent communication and interpersonal skills, able to work effectively across all levels of the organisation * Solid understanding of datacenter technologies (servers, networking, storage, virtualisation) within an operational support context * Strong Linux systems engineering experience, with proven troubleshooting across compute, storage, and network layers in production * Experience operating and debugging Kubernetes environments and distributed systems * Strong networking fundamentals (L2/L3, routing, VLANs, load balancing), with awareness of high-performance fabrics (RDMA/NVLink) * Experience with observability, monitoring, and incident response, including driving issues to resolution and contributing to post-incident improvements * Familiarity with SRE practices, including runbooks, process improvement, and reducing manual intervention * Experience with scripting and automation (Bash, Python, or similar) and Infrastructure as Code tools (e.g. Ansible, Terraform) * Strong analytical and problem-solving skills, with the ability to perform deep-dive investigations and root cause analysis * Familiarity with cloud infrastructure and virtualisation technologies; OpenStack experience preferred * Understanding of ITIL processes (incident, problem, and change management) Nice to Have: * Experience with GPU platforms (NVIDIA/AMD) and performance diagnostics (e.g. nvidia-smi, NCCL) * Exposure to HPC or distributed workloads (e.g. RDMA, InfiniBand, MPI) * Experience with CI/CD or GitOps tooling * Experience working in multi-region environments ## Description We are looking for an Infrastructure Support Lead, US to own the day-to-day operations of the US Support team. * You will ensure the team is performing, developing, and delivering exceptional service to Nscale's customers. * Reporting into the Infrastructure Support Manager, you will act as the bridge between strategic direction and frontline execution - translating priorities into action, managing team workload and productivity, and working alongside Senior Engineers to drive technical excellence. * You will own people management for all Engineers and Analysts in your region, while collaborating with Senior Engineers as peers on complex incidents and improvements. * You are a technically credible leader who is equally comfortable managing people as you are rolling up your sleeves on a complex incident. You operate with a bias for action, hold yourself and your team to high standards, and are driven to continuously improve both the service and the people delivering it., * Manage, coach, and mentor team members to deliver high-quality support and right-first-time resolution * Conduct regular 1:1s, performance reviews, and development planning * Set and monitor individual and team objectives, driving accountability and continuous improvement * Manage shift planning, rota coverage, and on-call scheduling * Identify skills gaps and drive upskilling through training, mentoring, and knowledge sharing across teams * Ensure roles, responsibilities, and expectations are clearly understood and consistently applied Ticket & Service Management * Own ticket queue management, ensuring accurate prioritisation and timely resolution * Monitor team productivity and workload trends, addressing bottlenecks to maintain service levels * Ensure adherence to ITIL processes across incidents, requests, changes, and problem management * Maintain accurate reporting on ticket status, SLA adherence, and team performance for the Infrastructure Support Manager Operational Excellence & Continuous Improvement * Identify regional risks and escalate priorities to the Infrastructure Support Manager where required * Improve dashboards, alerting, and runbooks to reduce repeat incidents and drive self-service resolution * Maintain consistent standards, processes, and documentation across the regional teams * Ensure compliance with audit, security, and operational requirements Technical Contribution & Support * Work alongside Senior Engineers on complex incidents, technical improvements, and operational tooling * Provide hands-on support across compute, storage, and networking layers * Support Kubernetes environments and Linux-based infrastructure at scale * Contribute to scripting and automation to improve operational workflows * Travel to Nscale or customer sites when needed to provide technical support Incident Management & Stakeholder Interaction * Act as the escalation point for complex or high-impact incidents including on-call (regional) * Lead post-incident reviews, identify recurring patterns, and ensure follow-up actions are tracked and delivered * Contribute to readiness and support planning for new services, deployments, and projects ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)