> Markdown version of [/jobs/ext/2710184-staff-infrastructure-software-engineer-fleet-automation](https://www.wearedevelopers.com/jobs/ext/2710184-staff-infrastructure-software-engineer-fleet-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Infrastructure Software Engineer, Fleet & Automation - **Company:** NuScale Power Corporation - **Location:** New York, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Intelligent Platform Management Interface, C++ (Programming Language), Cloud Computing, Code Review, Nvidia CUDA, Computer Engineering, Network Congestion, Continuous Integration, Data Centers, Data Center Infrastructure Management (CIM), Software Debugging, Linux, Distributed Systems, Ethernet, Firmware, InfiniBand, Networking Hardware, Python (Programming Language), Network Architecture, Networking Basics, Routing, OpenStack, Remote Direct Memory Access, Ansible, Prometheus, Workflow Management Systems, AI Infrastructure, Pulumi, Graphics Processing Unit (GPU), Cloud Platform System, Grafana, Event Driven Architecture, Kubernetes, Information Technology, Bare Metal, Slurm, Api Design, Terraform, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-infrastructure-software-engineer-fleet-automation-nscale-9007632 ## About the Role * 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems. * A Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience. * A strong software-engineering foundation in Python and/or Go, Java, C++, or similar languages, including API design, testing, code review, and production debugging. * Deep understanding of Linux, distributed systems, networking fundamentals, and systems performance; you are comfortable working across stateful and stateless services. * Experience designing and operating reliable automation or control-plane systems for complex infrastructure, large fleets, cloud platforms, or hardware lifecycle management. * Proven ability to take ambiguous technical problems from architecture through implementation and production operation, while influencing peers and stakeholders without relying on formal authority. * Hands-on experience with observability, monitoring, metrics, logs, tracing, alerting, incident response, capacity planning, and performance analysis. * Strong communication skills and sound technical judgment. You can explain trade-offs clearly, build alignment, and move work forward in a fast-changing environment. * A curious, pragmatic, high-ownership mindset. You enjoy finding the underlying cause of difficult problems and building the simplest durable solution. Strong Candidates Will Have * Direct experience with AI, GPU, HPC, or large-scale cloud infrastructure, including NVIDIA GPUs, CUDA, NVLink/NVSwitch, NCCL, and workload schedulers such as Slurm and Kubernetes. * Experience with high-performance datacentre networking, including InfiniBand, RoCE, Ethernet fabrics, routing, congestion control, topology-aware systems, or GPU Direct RDMA. * Experience with bare-metal lifecycle automation and infrastructure management tools such as Redfish, IPMI, PXE, MAAS, Ironic, NetBox, DCIM, OpenStack, or equivalent systems. * Experience with workflow orchestration and reliable automation systems such as Temporal, Airflow, Prefect, or event-driven architectures. * Experience with Kubernetes, containers, infrastructure as code (Terraform, Pulumi, Ansible), and public-cloud or private-cloud platforms. * Experience with observability platforms and high-cardinality telemetry, such as Prometheus, Grafana, OpenTelemetry, ELK, or equivalent. * Experience with hardware qualification, burn-in, validation, fleet health, or automated remediation for servers, GPUs, or network equipment. * A track record of technical leadership: setting direction, defining reusable patterns, developing other engineers, and improving the effectiveness of multiple teams. ## Description We're hiring a Staff Software Engineer to build the software, automation, and control-plane capabilities that manage Nscale's fleet of AI infrastructure at scale. Your work will improve the acceptance, performance, and scalability of our AI and high-performance computing environments - driving higher availability, faster capacity delivery, and lower operational load as Nscale grows into one of the world's leading neo-cloud providers. This is a senior individual-contributor role for an engineer who enjoys solving hard infrastructure problems at the intersection of software, GPUs, networking, and large-scale operations. You will have the autonomy to investigate problems, learn quickly, innovate, and deliver improvements wherever they create meaningful impact for the team and the platform. You will work closely with teams across Nscale - including Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware engineering - to translate operational challenges into reliable, scalable software. You will not need to own every component to make a difference: strong engineers identify opportunities, build a compelling case for a solution, and work with the right partners to deliver it. NOTE: We are hiring for various senior experience levels. The final leveling for the role will be based on your overall work experience, experience in AI Infra domain and interview feedback. What You'll Be Doing * Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems, balancing scalability, reliability, and maintainability. * Build and operate production-grade software, services, APIs, and automation that manage the lifecycle of GPU compute and supporting network infrastructure. * Own end-to-end workflows for fleet inventory, provisioning, configuration, hardware and firmware lifecycle management, validation, health monitoring, remediation, capacity, and reliability at scale. * Investigate complex production issues across hardware, GPUs, operating systems, networks, schedulers, and services; turn findings into durable software improvements rather than recurring manual work. * Build safe, observable, and auditable control-plane workflows that give operators clear visibility and reliable ways to act. * Establish engineering standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness. Use SLOs, telemetry, alerting, and postmortems to drive continuous improvement. * Partner with Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware teams to translate operational needs into robust, scalable automation. * Influence the evolution of adjacent systems and services through sound technical judgment, clear communication, and practical solutions. * Assess the impact of new hardware programmes on the software stack and ensure fleet-management capabilities are ready to support them. * Lead technical design reviews and incident deep-dives; mentor other engineers and raise the engineering bar across the organization. * Use AI-assisted development tools to increase delivery leverage while maintaining a high bar for correctness, security, and operational safety., The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)