Staff Infrastructure Software Engineer, Fleet & Automation

NuScale Power Corporation
New York, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Airflow Intelligent Platform Management Interface C++ (Programming Language) Cloud Computing Code Review Nvidia CUDA Computer Engineering Network Congestion Continuous Integration
+31 more
Data Centers Data Center Infrastructure Management (CIM) Software Debugging Linux Distributed Systems Ethernet Firmware InfiniBand Networking Hardware Python (Programming Language) Network Architecture Networking Basics Routing OpenStack Remote Direct Memory Access Ansible Prometheus Workflow Management Systems AI Infrastructure Pulumi Graphics Processing Unit (GPU) Cloud Platform System Grafana Event Driven Architecture Kubernetes Information Technology Bare Metal Slurm Api Design Terraform Golang

Job description

We’re hiring a Staff Software Engineer to build the software, automation, and control-plane capabilities that manage Nscale’s fleet of AI infrastructure at scale. Your work will improve the acceptance, performance, and scalability of our AI and high-performance computing environments - driving higher availability, faster capacity delivery, and lower operational load as Nscale grows into one of the world’s leading neo-cloud providers.

This is a senior individual-contributor role for an engineer who enjoys solving hard infrastructure problems at the intersection of software, GPUs, networking, and large-scale operations. You will have the autonomy to investigate problems, learn quickly, innovate, and deliver improvements wherever they create meaningful impact for the team and the platform.

You will work closely with teams across Nscale - including Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware engineering - to translate operational challenges into reliable, scalable software. You will not need to own every component to make a difference: strong engineers identify opportunities, build a compelling case for a solution, and work with the right partners to deliver it.

NOTE: We are hiring for various senior experience levels. The final leveling for the role will be based on your overall work experience, experience in AI Infra domain and interview feedback.

What You’ll Be Doing

  • Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems, balancing scalability, reliability, and maintainability.
  • Build and operate production-grade software, services, APIs, and automation that manage the lifecycle of GPU compute and supporting network infrastructure.
  • Own end-to-end workflows for fleet inventory, provisioning, configuration, hardware and firmware lifecycle management, validation, health monitoring, remediation, capacity, and reliability at scale.
  • Investigate complex production issues across hardware, GPUs, operating systems, networks, schedulers, and services; turn findings into durable software improvements rather than recurring manual work.
  • Build safe, observable, and auditable control-plane workflows that give operators clear visibility and reliable ways to act.
  • Establish engineering standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness. Use SLOs, telemetry, alerting, and postmortems to drive continuous improvement.
  • Partner with Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware teams to translate operational needs into robust, scalable automation.
  • Influence the evolution of adjacent systems and services through sound technical judgment, clear communication, and practical solutions.
  • Assess the impact of new hardware programmes on the software stack and ensure fleet-management capabilities are ready to support them.
  • Lead technical design reviews and incident deep-dives; mentor other engineers and raise the engineering bar across the organization.
  • Use AI-assisted development tools to increase delivery leverage while maintaining a high bar for correctness, security, and operational safety., The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

Requirements

  • 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems.
  • A Bachelor’s degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
  • A strong software-engineering foundation in Python and/or Go, Java, C++, or similar languages, including API design, testing, code review, and production debugging.
  • Deep understanding of Linux, distributed systems, networking fundamentals, and systems performance; you are comfortable working across stateful and stateless services.
  • Experience designing and operating reliable automation or control-plane systems for complex infrastructure, large fleets, cloud platforms, or hardware lifecycle management.
  • Proven ability to take ambiguous technical problems from architecture through implementation and production operation, while influencing peers and stakeholders without relying on formal authority.
  • Hands-on experience with observability, monitoring, metrics, logs, tracing, alerting, incident response, capacity planning, and performance analysis.
  • Strong communication skills and sound technical judgment. You can explain trade-offs clearly, build alignment, and move work forward in a fast-changing environment.
  • A curious, pragmatic, high-ownership mindset. You enjoy finding the underlying cause of difficult problems and building the simplest durable solution.

Strong Candidates Will Have

  • Direct experience with AI, GPU, HPC, or large-scale cloud infrastructure, including NVIDIA GPUs, CUDA, NVLink/NVSwitch, NCCL, and workload schedulers such as Slurm and Kubernetes.
  • Experience with high-performance datacentre networking, including InfiniBand, RoCE, Ethernet fabrics, routing, congestion control, topology-aware systems, or GPU Direct RDMA.
  • Experience with bare-metal lifecycle automation and infrastructure management tools such as Redfish, IPMI, PXE, MAAS, Ironic, NetBox, DCIM, OpenStack, or equivalent systems.
  • Experience with workflow orchestration and reliable automation systems such as Temporal, Airflow, Prefect, or event-driven architectures.
  • Experience with Kubernetes, containers, infrastructure as code (Terraform, Pulumi, Ansible), and public-cloud or private-cloud platforms.
  • Experience with observability platforms and high-cardinality telemetry, such as Prometheus, Grafana, OpenTelemetry, ELK, or equivalent.
  • Experience with hardware qualification, burn-in, validation, fleet health, or automated remediation for servers, GPUs, or network equipment.
  • A track record of technical leadership: setting direction, defining reusable patterns, developing other engineers, and improving the effectiveness of multiple teams.

Benefits & conditions

  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.

About the company

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all