> Markdown version of [/jobs/ext/2074268-software-engineer-fleet-automation](https://www.wearedevelopers.com/jobs/ext/2074268-software-engineer-fleet-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Fleet Automation - **Company:** NorthMark Strategies - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), C Sharp (Programming Language), Ubuntu (Operating System), Software Debugging, Linux, Enterprise Messaging Systems, NoSQL, Red Hat Enterprise Linux, Prometheus, Software Engineering, TypeScript, Grafana, Hardware Testing, Backend, Event Driven Architecture, Build Management, Infrastructure Automation Frameworks, Information Technology, Bare Metal, Performance Monitor, Apache Kafka, Golang - **Published:** August 15, 2026 - **Apply:** https://www.dice.com/job-detail/b9a39d80-1892-4110-9fdc-9964120dd79e ## About the Role * Bachelor's Degree in Computer Science, Software Engineering, or equivalent practical experience. * 5+ years of software engineering experience building production backend services or infrastructure automation tooling. * Proficiency in Go, C#, or TypeScript. * Experience designing and working with relational and NoSQL databases to support stateful automation workflows and internal platform services. * Solid understanding of Linux systems - networking, storage, process management, and debugging on Ubuntu or RHEL variants. * Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, Alertmanager, ELK) in a production environment. * Familiarity with GPU compute infrastructure and NVIDIA tooling (DCGM, nvidia-smi, NVIDIA Container Toolkit) is a strong plus. * Exposure to event-driven architectures or messaging platforms (e.g. Kafka) is a plus for teams building automation workflows across distributed services. * Strong communication skills and a collaborative mindset - comfortable navigating ambiguity, taking initiative, and working across Infrastructure, Operations, and Research teams. It is impossible to list every requirement for, or responsibility of, any position. Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company's needs may change over time. Therefore, the above job description is not comprehensive or exhaustive. The Company reserves the right to adjust, add to or eliminate any aspect of the above description. The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs. Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future. ## Description NMC is seeking a Software Engineer to join the Fleet Automation team within the HPC & Infrastructure organization. This team owns the systems and tooling that keep hundreds of high-performance GPU compute nodes provisioned, configured, and operating at peak efficiency - spanning bare-metal provisioning, lifecycle management, and automated remediation at scale. In this role, you will design and build the automation platforms, internal services, and APIs that allow NMC to operate its growing fleet with speed and reliability. You will work at the intersection of software engineering and infrastructure - writing production-quality Go, C#, and TypeScript services that directly manage physical hardware, integrate with orchestration layers, and surface actionable observability to operations and on-call teams. You will collaborate closely with Infrastructure Engineers, Network Engineers, and Research & Client teams to translate operational pain points into durable, maintainable automation. You will participate in on-call rotations and take ownership of system health through proactive monitoring, alerting, and incident response. The ideal candidate is equally comfortable architecting backend services and debugging Linux systems, thrives in ambiguous environments, and takes pride in eliminating manual toil through well-crafted tooling. This role is based in Dallas, TX out of our Victory Commons office., * Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes. * Develop APIs and service integrations that enable Infrastructure and Operations teams to deploy, image, validate, and decommission hardware with minimal manual intervention. * Build and maintain backend services in Go, C#, and TypeScript with a strong focus on reliability, testability, and long-term maintainability. * Design and evolve data models and persistent state for automation workflows, working across relational and NoSQL databases as appropriate. * Build and maintain CI/CD pipelines that gate configuration changes, run automated hardware validation tests, and promote changes safely across environments. * Instrument systems for observability - designing metrics, alerts, and dashboards in Prometheus and Grafana that provide real-time fleet health visibility to on-call teams. * Participate in on-call rotations; own incident response, post-mortems, and follow-through on reliability improvements across the fleet. * Identify systemic gaps in fleet reliability and efficiency and champion engineering solutions that reduce operational toil at scale. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025)