GPU Infrastructure Engineer
Techclub, Inc
Fort Worth, TX, United States
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.disabledperson.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Systems Engineering
Bash Shell
Ubuntu (Operating System)
Nvidia CUDA
Computer Programming
Data Centers
Linux
Hyper-V
Python (Programming Language)
Linux Kernel
Linux System Administration
+15 more
Windows Servers
Nagios
PCI Express
Performance Tuning
Red Hat Enterprise Linux
Ansible
Virtualization Technology
Diagnostic Tools
Computer Networking Systems
Grafana
Firewalls (Computer Science)
Infrastructure Automation Frameworks
Influxdb
Hardware Infrastructure
Vmware
Job description
- Design, deploy, maintain, and troubleshoot enterprise-scale Linux and Windows server infrastructure.
- Support hyperscale GPU and compute environments used for AI training and inference workloads.
- Troubleshoot NVIDIA and AMD GPU platforms, including GPU memory errors, driver issues, CUDA/ROCm failures, PCIe problems, and system-level hardware faults.
- Work with GPU technologies and related OCP-based hardware to optimize system performance.
- Diagnose server, rack, power, networking, PCIe, GPU, NIC, and operating-system issues.
- Utilize Linux kernel tools, NVIDIA/AMD diagnostic utilities, system logs, and custom troubleshooting tools to identify and resolve infrastructure problems.
- Develop automation and infrastructure tooling using Python, Bash, Ansible, Chef, and related technologies.
- Improve operational tools and automation used across large server fleets.
- Support VMware and Hyper-V virtualization environments where required.
- Administer and troubleshoot networking infrastructure, including switches, firewalls, and data center connectivity.
- Monitor infrastructure health and performance using tools such as Grafana, InfluxDB, Telegraf, Nagios, and related monitoring platforms.
- Participate in on-call rotations and provide critical incident response for production infrastructure.
- Identify systemic infrastructure issues and develop scalable solutions to prevent recurrence.
- Collaborate with engineering and infrastructure teams to improve reliability, availability, and operational efficiency.
Requirements
- 10+ years of experience in systems, infrastructure, production engineering, or related roles.
- Strong Linux administration and troubleshooting experience, preferably with RHEL and Ubuntu.
- Hands-on experience with enterprise server and data center infrastructure.
- Experience troubleshooting GPU-based infrastructure and high-performance compute environments.
- Strong understanding of server hardware, PCIe, networking, storage, and operating-system diagnostics.
- Programming/scripting experience with Python and Bash.
- Experience with infrastructure automation tools such as Ansible or Chef.
- Strong troubleshooting and root-cause-analysis skills.
- Experience working in large-scale production environments.
- Excellent communication and cross-functional collaboration skills., The ideal candidate is a hands-on infrastructure professional who can operate at both the hardware and software layers, from diagnosing GPU/server failures and Linux kernel issues to developing automation that improves reliability across large-scale infrastructure. Experience supporting hyperscale AI/GPU environments, combined with strong systems engineering and production troubleshooting skills, is highly valued.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.disabledperson.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
CS
Christina Schaireiter
3 months ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
19 days ago
LM
Luis Minvielle
A Guide to Green Tech and Green IT Careers
over 2 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago