> Markdown version of [/jobs/ext/2713196-staff-software-engineer-fleet-management](https://www.wearedevelopers.com/jobs/ext/2713196-staff-software-engineer-fleet-management). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - Fleet Management - **Company:** NuScale Power Corporation - **Location:** United States - **Salary:** $220,000.0 - $320,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Intelligent Platform Management Interface, Cursor (Graphical User Interface Elements), Dynamic Host Configuration Protocol, Data Center Infrastructure Management (CIM), Distributed Systems, InfiniBand, Python (Programming Language), Network Configuration and Change Management, OpenStack, Performance Tuning, Systems Integration, Workflow Management Systems, Data Logging, Network Switches, Pulumi, State Machines, Build Management, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Free and Open-Source Software, Hardware Infrastructure, Terraform, Open Network Automation Platform, Dynatrace - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-software-engineer-fleet-management-nscale-8958557 ## About the Role * Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering. * Strong proficiency in Python - Fleet Manager is built entirely in Python. * Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling. * Track record of delivering automation systems from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization). * Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority. * You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement. * You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery. * Excellent communication skills to build consensus with stakeholders, both internally and externally, in a fast-paced, high-agency environment. Preferred * Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar * Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems * Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation * Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows * GPU infrastructure experience: health monitoring, burn-in testing, or cluster management * HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE) * Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP * Open-source contributions in infrastructure automation or cloud-native tooling ## Description Nscale is hiring a Staff Software Engineer to build Fleet Manager - the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale. This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high - the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand. This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on. What you'll work on * Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready - BMC configuration, DHCP reservations, and provisioning state machines. * Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet. * Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other. * GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production. * Network configuration: switch lifecycle automation and network state management across the fleet. * Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure. * Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively., * Domain-level technical direction. Own the architecture for a major Fleet Manager domain - such as provisioning, validation, or remediation - influencing engineers across the team and adjacent squads. * Design and build production-grade automation. Implement device provisioning, burn-in testing, network configuration, and hardware health validation workflows in Python. * Engineer for reliability and auditability. Treat idempotency, resumability, checkpointing, retries, replay, and failure handling as first-class design concerns. * Integrate broadly. Connect Fleet Manager with datacenter infrastructure management systems, cloud orchestration platforms, and bare-metal provisioning tools. * Create leverage through standards. Establish shared patterns, libraries, conventions, and operational runbooks that other engineers build on. * Own production outcomes. Operate what you build with strong observability, alerting, incident response, and day-2 operational discipline. * Mentor and influence. Raise the bar through design reviews, implementation guidance, and operational best practices. * Use AI to accelerate delivery while maintaining architectural coherence. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)