> Markdown version of [/jobs/ext/2098354-staff-principal-deployment-automation-engineer](https://www.wearedevelopers.com/jobs/ext/2098354-staff-principal-deployment-automation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff/Principal Deployment Automation Engineer - **Company:** Crusoe's Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $250,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Cloud Computing, Computer Clusters, Configuration Management, Nvidia CUDA, Continuous Integration, Software Debugging, Linux, Memory Management, InfiniBand, Python (Programming Language), PostgreSQL, Node.Js, PCI Express, Remote Direct Memory Access, Ansible, Graphics Processing Unit (GPU), Computer Network Technologies, Delivery Pipeline, Saltstack, Gitlab, Integration Tests, Kubernetes, Information Technology, Deployment Automation, Bare Metal, Puppet, Terraform, Docker - **Published:** August 17, 2026 - **Apply:** https://www.dice.com/job-detail/77d021bd-99f1-4d7e-a320-4681614c5453 ## About the Role * Education & Experience: 12+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field. * Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes. * Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres. * CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters. * Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack. * Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios. * System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU). * Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context. * Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems. Bonus Points: * Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures. * Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf). * Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins). ## Description As a Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters. You will own the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud. Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters., * Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe's AI Cloud Stack. * CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications. * Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads. * Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc. * Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary. * Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments. * Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)