> Markdown version of [/jobs/ext/3617580-staff-software-engineer-principal-software-engineer-infrastructure-engineer-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3617580-staff-software-engineer-principal-software-engineer-infrastructure-engineer-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer / Principal Software Engineer / Infrastructure Engineer / Platform Engineer - **Company:** USA, UTILITIES SERVICES ALLIANCE, INC. - **Location:** United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Intelligent Platform Management Interface, Cloud Computing, Data Centers, Cursor, Data Center Infrastructure Management (CIM), Programming Tools, Distributed Systems, Monitoring of Systems, Python (Programming Language), Network Architecture, Ansible, Systems Integration, Workflow Management Systems, AI Infrastructure, State Machines, Event Driven Architecture, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Hardware Infrastructure, Claude, Terraform, Open Network Automation Platform - **Published:** October 8, 2026 - **Apply:** https://www.thejobnetwork.com/job/ade48e40-abe0-43cf-b3e0-038281247f09/staff-software-engineer ## About the Role The ideal candidate will have extensive experience building distributed systems, infrastructure automation and production platforms, with strong Python skills and experience managing large-scale compute, datacenter or hardware infrastructure., * Extensive experience building and operating distributed systems in production \n * Strong Python software engineering skills \n * Experience with workflow orchestration, state machines or event-driven architectures \n * Experience building reliable automation for infrastructure provisioning and lifecycle management \n * Strong understanding of idempotency, retries, checkpoints, recovery and failure handling \n * Experience with bare-metal infrastructure, server provisioning or hardware lifecycle management \n * Experience with BMC, IPMI, Redfish, PXE, MAAS, Ironic or NetBox is highly desirable \n * Experience with GPU infrastructure, AI clusters, HPC or datacenter environments is advantageous \n * Strong Kubernetes, Terraform, Ansible or cloud infrastructure experience \n * Experience with networking and network automation, particularly at datacenter scale \n * Experience building health monitoring, validation, observability and automated remediation systems \n * Experience integrating infrastructure platforms, APIs, monitoring systems, credential stores and DCIM tools \n * Strong understanding of production reliability, scalability and security \n * Staff/Principal-level technical leadership with experience influencing architecture across teams \n * Experience using modern AI-assisted development tools such as Claude, Cursor or similar is desirable \n ## Description The role will involve building and owning a Fleet Management platform responsible for provisioning, testing, monitoring and remediating GPU servers and network infrastructure at scale., The platform will need to support thousands of concurrent workflows, with strong guarantees around reliability, auditability, idempotency, resumability and failure recovery. You will work closely with infrastructure, hardware, networking and operations teams to build highly scalable systems powering next-generation AI infrastructure.