Site Reliability Engineer (SRE) GPU Infrastructure

Neuveon Inc
Santa Clara, United States
6 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computer Programming Computer Engineering Data Centers Linux Firmware Python (Programming Language) Linux Kernel Linux System Administration AI Infrastructure Graphics Processing Unit (GPU) Infrastructure Automation Frameworks
+3 more
Information Technology Hardware Infrastructure Hardware Debugging

Job description

  • Troubleshoot complex Linux, GPU, server hardware, firmware, and infrastructure issues.
  • Analyze system/kernel logs and BMC/Redfish telemetry to identify root causes.
  • Support hardware provisioning, repair, deployment, and production validation.
  • Develop automation, diagnostics, provisioning, and hardware repair tools.
  • Test and validate next-generation AI servers and GPU platforms.
  • Develop operational documentation, procedures, and best practices.
  • Collaborate with Hardware Engineering, Data Center Operations, and Capacity Planning teams.
  • Provide on-call and remote operational support as required.

Requirements

  • Strong Linux administration and Linux internals knowledge.
  • Proven experience troubleshooting server hardware and infrastructure.
  • Strong understanding of GPU, hardware, firmware, networking, and systems troubleshooting.
  • Excellent root-cause analysis and problem-solving skills.
  • Experience with infrastructure provisioning and operations.
  • Strong communication and cross-functional collaboration skills.
  • Bachelor’s degree in Computer Science or related field, or equivalent experience.

Nice to Have

  • Experience with large-scale GPU/AI infrastructure.
  • Programming experience with Python, Go, or similar languages.
  • Experience developing infrastructure or hardware automation tools.
  • Experience with BMC/Redfish and data center hardware operations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

48 sec

Building secure data environments with dedicated GPU infrastructure

Hissan Usmani Hissan Usmani · World Congress 2025

Videos

See all

Related articles

See all