> Markdown version of [/jobs/ext/2658620-software-engineer-ii-dcie](https://www.wearedevelopers.com/jobs/ext/2658620-software-engineer-ii-dcie). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer II (DCIE) - **Company:** Crusoe's Inc - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $140,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Automation of Tests, Cloud Computing, Data Centers, Distributed Systems, Python (Programming Language), Software Engineering, Rust (Programming Language), Pytorch, Kubernetes, Golang, Programming Languages - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=00a23f0217a5e815 ## About the Role * 2-3 years of software engineering experience. * The ability to identify a problem, rapidly develop a scalable solution and ship it. * Ability to lean in and assist team members working on critical or complex technical initiatives. * Ability to set the technical direction for a specific project and execute. * Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.) * Strength in at least one programming language - Go, Python, Java, Rust. * Strong analytical and problem-solving skills. * Excellent communication and collaboration skills. * Ability to work independently and within a team, * Experience with Temporal and Kubernetes. * Experience working directly with hardware vendors. * Background in large-scale GPU fleet operations or hyperscale data center environments. ## Description We are seeking a highly skilled and motivated Software Engineer to join Crusoe's Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters. The ideal new team member will be a hands-on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe's rapidly growing GPU fleet., * Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems. * Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X. * Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware. * In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment. * Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance. * Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success. * Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [How we will build the software of tomorrow](https://www.wearedevelopers.com/videos/522-how-we-will-build-the-software-of-tomorrow) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Is Software Engineering Hard?](https://www.wearedevelopers.com/magazine/448-is-software-engineering-hard) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)