Site Reliability Engineer (SRE) GPU Infrastructure
Neuveon Inc
Santa Clara, United States
6 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Computer Programming
Computer Engineering
Data Centers
Linux
Firmware
Python (Programming Language)
Linux Kernel
Linux System Administration
AI Infrastructure
Graphics Processing Unit (GPU)
Infrastructure Automation Frameworks
+3 more
Information Technology
Hardware Infrastructure
Hardware Debugging
Job description
- Troubleshoot complex Linux, GPU, server hardware, firmware, and infrastructure issues.
- Analyze system/kernel logs and BMC/Redfish telemetry to identify root causes.
- Support hardware provisioning, repair, deployment, and production validation.
- Develop automation, diagnostics, provisioning, and hardware repair tools.
- Test and validate next-generation AI servers and GPU platforms.
- Develop operational documentation, procedures, and best practices.
- Collaborate with Hardware Engineering, Data Center Operations, and Capacity Planning teams.
- Provide on-call and remote operational support as required.
Requirements
- Strong Linux administration and Linux internals knowledge.
- Proven experience troubleshooting server hardware and infrastructure.
- Strong understanding of GPU, hardware, firmware, networking, and systems troubleshooting.
- Excellent root-cause analysis and problem-solving skills.
- Experience with infrastructure provisioning and operations.
- Strong communication and cross-functional collaboration skills.
- Bachelor’s degree in Computer Science or related field, or equivalent experience.
Nice to Have
- Experience with large-scale GPU/AI infrastructure.
- Programming experience with Python, Go, or similar languages.
- Experience developing infrastructure or hardware automation tools.
- Experience with BMC/Redfish and data center hardware operations.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
3 months ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
24 days ago
LM
Luis Minvielle
6 Emerging Technologies We’ll Learn About in 2025
over 2 years ago