> Markdown version of [/jobs/ext/2606399-gpu-infrastructure-noc-engineer](https://www.wearedevelopers.com/jobs/ext/2606399-gpu-infrastructure-noc-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU Infrastructure NOC Engineer - **Company:** Orion Placement - **Location:** Pittsburgh, PA, United States (Remote available) - **Experience:** Experienced - **Salary:** $75,000.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Data Centers, Monitoring of Systems, Python (Programming Language), Nagios, Reliability Engineering, Software Deployment, AI Infrastructure, Datadog, Scripting, Computer Network Operations, High Performance Computing, Grafana, Software Troubleshooting, Hardware Infrastructure, 3-tier Architectures, Pagerduty - **Published:** August 31, 2026 - **Apply:** https://www.careerjet.com/jobad/useb3dce87aec3ffd60cd6832734511944 ## About the Role * 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment. * Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments. * Hands-on Python or Bash scripting experience. * Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms. * Strong troubleshooting and incident-response skills. * Experience working with network, compute, storage, or data center infrastructure. * Ability to understand technical issues quickly and communicate effectively during incidents. * Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement. * Must be comfortable working rotating 24/7 shifts, including nights and weekends. * Ability to work independently in a remote environment while collaborating effectively with global technical teams. * Additional languages beyond English are a plus., Requirements: Must-have: remote nationwide. 2+ yrs NOC, network ops, or infrastructure monitoring exp, Must accept 24/7 rotating shifts, including nights/weekends (Y/N hard deal break), must have GPU/HPC exp. Must have Python/Bash exp Required Years as Associate: 2+, Monitor and support live GPU/HPC clusters in a 24/7 NOC while handling incidents, SLA performance, automation, and runbook improvements. Must have 2+ years NOC/network/infrastructure monitoring, GPU/HPC experience, and Python or Bash. Must accept rotating nights/weekends., Experience: 2+ Good fit job titles/keywords for candidates: NOC Engineer, NOC Analyst, Network Operations Engineer, Infrastructure Operations Engineer, GPU NOC, HPC Operations, GPU Infrastructure Engineer, AI Infrastructure Engineer, Systems Operations, SRE, Data Center Operations, Python, Bash # of hires needed: 1 ## Description * Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads. * Go beyond traditional NOC monitoring by building automation, improving tooling, and making operations smarter. * Get hands-on exposure to GPU, HPC, networking, data center infrastructure, and AI-enabled operations. * Play a direct role in protecting customer uptime and meeting demanding SLA commitments. * Help create and continuously improve runbooks, SOPs, monitoring processes, and incident-response workflows. * Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems. * Join a fast-growing environment where your ideas can directly improve how the NOC operates. * Receive bonus and equity opportunities in addition to competitive base compensation., Location: Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts, including nights and weekends. Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts, including nights and weekends. These are hard requirements. About Us We are building next-generation AI infrastructure designed to give enterprise customers fast, flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable, responsive, and continuously improving. Confidential Employer., * Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments. * Triage, troubleshoot, and resolve incidents while maintaining SLA requirements. * Identify infrastructure issues early and take proactive action before they become customer-impacting incidents. * Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners. * Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work. * Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows. * Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency. * Create, maintain, and continuously improve technical runbooks and standard operating procedures. * Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements. * Track SLA and incident metrics and identify opportunities to improve reliability and response times. * Communicate clearly and proactively with customers and internal stakeholders during incidents. * Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues. * Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)