> Markdown version of [/jobs/ext/2694785-software-engineer-fleet-automation-gpu-infrastructure](https://www.wearedevelopers.com/jobs/ext/2694785-software-engineer-fleet-automation-gpu-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Fleet Automation / GPU Infrastructure - **Company:** GTN Technical Staffing - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Systems Engineering, C Sharp (Programming Language), Ubuntu (Operating System), Continuous Delivery, Continuous Integration, Relational Databases, Linux, Distributed Systems, Automation of Marketing, Enterprise Messaging Systems, NoSQL, Red Hat Enterprise Linux, Reliability Engineering, Prometheus, Software Engineering, TypeScript, Graphics Processing Unit (GPU), Grafana, Hardware Testing, Software Troubleshooting, Backend, Event Driven Architecture, Build Management, Information Technology, Bare Metal, Apache Kafka, Hardware Infrastructure - **Published:** September 3, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-software-engineer-fleet-automation-gpu-infrastructure-dallas-tx--8c12079d-500d-4b18-833b-e0e605f6ee74 ## About the Role * 5+ years of software engineering experience building production backend services, infrastructure platforms, or automation tooling. * Strong development experience with at least one of the following: * + Go + C# + TypeScript * Experience designing APIs, backend services, and distributed or stateful systems. * Strong experience with relational and/or NoSQL databases. * Solid Linux systems knowledge, including: * + Networking + Storage + Process management + System troubleshooting + Ubuntu and/or RHEL environments * Experience building and maintaining CI/CD pipelines. * Hands-on experience with production monitoring and observability platforms such as: * + Prometheus + Grafana + Alertmanager + ELK * Strong troubleshooting and problem-solving skills across both software and infrastructure environments. Preferred Experience * Experience supporting GPU, HPC, AI/ML, or large-scale compute infrastructure. * Familiarity with NVIDIA technologies such as: * + DCGM + nvidia-smi + NVIDIA Container Toolkit * Experience with bare-metal provisioning and hardware lifecycle automation. * Exposure to event-driven architectures and messaging platforms such as Kafka. * Experience working with infrastructure, network, SRE, or platform engineering teams. * Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience., Apache Kafka, Application Programming Interface (API), Artificial Intelligence (AI), Automation, Automation Systems, CPU (Central Processing Unit), Computer Science, Computer Systems, Continuous Deployment/Delivery, Continuous Integration, Data Modeling, Distributed Computing, Engineering, GPU (Graphics Processing Unit), Hardware Quality Assurance, Identify Issues, Incident Response, Linux Operating System, Machine Tool, Microsoft C# (C Sharp), NoSQL, On Call, Operations Research, Problem Solving Skills, Process Management, Production Control, Red Hat Linux Operating System, Relational Databases (RDBMS), Reliability Engineering, Reporting Dashboards, Root Cause Analysis, Software Engineering, Statement of Work (SOW), Systems Engineering, Technical Recruiting, Ubuntu, Validation Testing, Vehicle Fleets ## Description Our client is seeking a Senior Software Engineer to join its Fleet Automation team supporting large-scale HPC and GPU infrastructure. This team builds the software, automation, and internal platforms used to provision, configure, monitor, and manage hundreds of high-performance GPU and CPU compute nodes. The environment sits at the intersection of software engineering, infrastructure, and distributed systems, with a strong focus on eliminating manual operational work through scalable automation. This is a hands-on engineering role for someone who enjoys building backend services while also understanding how those services interact with Linux systems, physical hardware, networking, storage, and GPU infrastructure. What You'll Do * Design and build fleet automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes. * Develop internal services and APIs that automate hardware deployment, imaging, remediation, and decommissioning. * Build reliable backend services using Go, C#, and/or TypeScript. * Design data models and persistent state for automation workflows using relational and NoSQL databases. * Develop and maintain CI/CD pipelines for infrastructure and configuration changes. * Automate hardware validation and testing across large-scale compute environments. * Build observability, monitoring, dashboards, and alerting using tools such as Prometheus, Grafana, Alertmanager, and ELK. * Partner closely with Infrastructure, Network, Operations, and Research teams to identify operational pain points and automate repeatable processes. * Participate in on-call rotations and support incident response, root-cause analysis, and post-incident reliability improvements. * Identify systemic infrastructure issues and develop engineering solutions that improve fleet reliability, efficiency, and scalability. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)