Staff Software Engineer / Principal Software Engineer / Infrastructure Engineer / Platform Engineer

USA, UTILITIES SERVICES ALLIANCE, INC.
United States
4 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Intelligent Platform Management Interface Cloud Computing Data Centers Cursor Data Center Infrastructure Management (CIM) Programming Tools Distributed Systems Monitoring of Systems Python (Programming Language) Network Architecture
+13 more
Ansible Systems Integration Workflow Management Systems AI Infrastructure State Machines Event Driven Architecture Kubernetes Infrastructure Automation Frameworks Bare Metal Hardware Infrastructure Claude Terraform Open Network Automation Platform

Job description

The role will involve building and owning a Fleet Management platform responsible for provisioning, testing, monitoring and remediating GPU servers and network infrastructure at scale., The platform will need to support thousands of concurrent workflows, with strong guarantees around reliability, auditability, idempotency, resumability and failure recovery. You will work closely with infrastructure, hardware, networking and operations teams to build highly scalable systems powering next-generation AI infrastructure.

Requirements

The ideal candidate will have extensive experience building distributed systems, infrastructure automation and production platforms, with strong Python skills and experience managing large-scale compute, datacenter or hardware infrastructure., * Extensive experience building and operating distributed systems in production \n

  • Strong Python software engineering skills \n

  • Experience with workflow orchestration, state machines or event-driven architectures \n

  • Experience building reliable automation for infrastructure provisioning and lifecycle management \n

  • Strong understanding of idempotency, retries, checkpoints, recovery and failure handling \n

  • Experience with bare-metal infrastructure, server provisioning or hardware lifecycle management \n

  • Experience with BMC, IPMI, Redfish, PXE, MAAS, Ironic or NetBox is highly desirable \n

  • Experience with GPU infrastructure, AI clusters, HPC or datacenter environments is advantageous \n

  • Strong Kubernetes, Terraform, Ansible or cloud infrastructure experience \n

  • Experience with networking and network automation, particularly at datacenter scale \n

  • Experience building health monitoring, validation, observability and automated remediation systems \n

  • Experience integrating infrastructure platforms, APIs, monitoring systems, credential stores and DCIM tools \n

  • Strong understanding of production reliability, scalability and security \n

  • Staff/Principal-level technical leadership with experience influencing architecture across teams \n

  • Experience using modern AI-assisted development tools such as Claude, Cursor or similar is desirable \n

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Loading talks and stories from around this role…