GPU Infrastructure NOC Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
- Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads.
- Go beyond traditional NOC monitoring by building automation, improving tooling, and making operations smarter.
- Get hands-on exposure to GPU, HPC, networking, data center infrastructure, and AI-enabled operations.
- Play a direct role in protecting customer uptime and meeting demanding SLA commitments.
- Help create and continuously improve runbooks, SOPs, monitoring processes, and incident-response workflows.
- Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems.
- Join a fast-growing environment where your ideas can directly improve how the NOC operates.
-
Receive bonus and equity opportunities in addition to competitive base compensation., Location: Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts, including nights and weekends. Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts, including nights and weekends. These are hard requirements. About Us We are building next-generation AI infrastructure designed to give enterprise customers fast, flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable, responsive, and continuously improving. Confidential Employer., * Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
- Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
- Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
- Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners.
- Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
- Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows.
- Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency.
- Create, maintain, and continuously improve technical runbooks and standard operating procedures.
- Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.
- Track SLA and incident metrics and identify opportunities to improve reliability and response times.
- Communicate clearly and proactively with customers and internal stakeholders during incidents.
- Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues.
- Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts.
Requirements
- 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
- Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
- Hands-on Python or Bash scripting experience.
- Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.
- Strong troubleshooting and incident-response skills.
- Experience working with network, compute, storage, or data center infrastructure.
- Ability to understand technical issues quickly and communicate effectively during incidents.
- Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement.
- Must be comfortable working rotating 24/7 shifts, including nights and weekends.
- Ability to work independently in a remote environment while collaborating effectively with global technical teams.
- Additional languages beyond English are a plus., Requirements: Must-have: remote nationwide. 2+ yrs NOC, network ops, or infrastructure monitoring exp, Must accept 24/7 rotating shifts, including nights/weekends (Y/N hard deal break), must have GPU/HPC exp. Must have Python/Bash exp Required Years as Associate: 2+, Monitor and support live GPU/HPC clusters in a 24/7 NOC while handling incidents, SLA performance, automation, and runbook improvements. Must have 2+ years NOC/network/infrastructure monitoring, GPU/HPC experience, and Python or Bash. Must accept rotating nights/weekends., Experience: 2+ Good fit job titles/keywords for candidates: NOC Engineer, NOC Analyst, Network Operations Engineer, Infrastructure Operations Engineer, GPU NOC, HPC Operations, GPU Infrastructure Engineer, AI Infrastructure Engineer, Systems Operations, SRE, Data Center Operations, Python, Bash # of hires needed: 1
Benefits & conditions
- Work directly with advanced GPU and HPC infrastructure rather than generic IT systems.
- Build technical depth across AI compute, networking, data centers, monitoring, and automation.
- Have a voice in how the NOC operates and help improve processes rather than simply following them.
- Work with modern monitoring, automation, and AI-enabled operations tools.
- Gain exposure to complex enterprise infrastructure and high-availability environments.
- Remote nationwide flexibility with multiple shift options.
- Medical, dental, and vision insurance.
- 401(k).
- Paid maternity and paternity leave.
- Bonus and equity opportunities., * Dental insurance
- Paid time off
- Retirement plan
- Vision insurance
About the company
Recruiter Submission: To submit, cancel - GPU Infrastructure NOC Engineer - Axe Compute - JPC-1916 - source New Job Order Alert
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Top 6 Hackathons for Developers in 2023
Highest Paying Tech Companies for Developers
A Guide to Green Tech and Green IT Careers