> Markdown version of [/jobs/ext/3122698-software-engineer-server-fleet-infrastructure](https://www.wearedevelopers.com/jobs/ext/3122698-software-engineer-server-fleet-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Server Fleet Infrastructure - **Company:** Coreweave Inc - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Salary:** $139,000.0 - $242,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Computer Engineering, Continuous Integration, Data Centers, Software Debugging, Linux, Distributed Systems, Github, Node.Js, Ansible, Systems Integration, Concurrency, Backend, Build Management, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Bare Metal, Puppet - **Published:** September 28, 2026 - **Apply:** https://www.juju.com/job/16_6f8958a3 ## About the Role * 5+ years of experience in software or infrastructure engineering building and operating production backend or infrastructure services. * Proficiency in Go for building networked services and APIs (gRPC and REST) in production environments. * Experience designing and implementing distributed systems that operate reliably at scale, including concurrency, failure handling, and resiliency patterns. * Strong understanding of Linux systems and how to debug issues across processes, networking, and storage. * Experience with Kubernetes or similar container orchestration platforms and their APIs (for example, interacting with custom resources, controllers, or operators). * Familiarity with CI/CD tooling (such as Argo, Flux, or GitHub Actions) to ship and operate services safely and frequently. * Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Preferred: * Experience designing and operating services that automate lifecycle management for large fleets of physical servers or other hardware. * Experience with infrastructure automation and configuration management tools (for example, Ansible, Puppet, Chef, or Salt). * Experience integrating with vendor APIs and internal systems to coordinate multi-step operational workflows. * Experience with multi-datacenter or regionally distributed systems, including thinking about failure domains, capacity, and rollout strategies. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You enjoy building backend APIs and distributed services that automate real-world infrastructure operations at scale. * You are curious about how large fleets of bare metal servers are provisioned, updated, and repaired across many data centers. * You are excited to collaborate with hardware, data center, and platform engineers to turn complex procedures into reliable automation. ## Description The Fleet Provisioning Automation (FPA) team is responsible for the automated provisioning and lifecycle management of CoreWeave's rapidly growing fleet of hardware nodes and node types. The team streamlines and coordinates node bring-up, hardware RMA, data center operations, and platform services into a cohesive, high-reliability engine of fleet management. This group sits at the intersection of hardware, data center operations, and platform engineering, building the software that keeps our global fleet healthy and ready for customer workloads., As a Senior Software Engineer on the Fleet Provisioning Automation team, you will design and build backend services and APIs that automate provisioning, configuration, and lifecycle operations for CoreWeave's globally distributed bare metal fleet. You will primarily work in Go to implement gRPC APIs that integrate with Kubernetes, vendor and internal services, and data center tooling. Your work will focus on turning complex, multi-step operational procedures into simple, safe, and auditable automation for internal users operating at hyperscale. In this role, you will: * Design and implement backend services and APIs (primarily gRPC in Go) that orchestrate provisioning, configuration, and lifecycle operations across CoreWeave's global server fleet. * Develop Kubernetes custom resource definitions (CRDs) to automate provisioning and lifecycle management of CoreWeave's entire server fleet. * Model and evolve RPC schemas and data contracts used by other engineering and operations teams to integrate with fleet provisioning workflows. * Build integrations with vendor and internal APIs to make hardware and data center processes robust, transparent, and easy to operate at scale. * Collaborate closely with hardware engineering, data center operations, and platform teams to design solutions to problems of scale for multi-site deployment and management of CoreWeave's global hardware fleet. * Create test plans, deployment automation, and supporting tooling that enable safe rollouts and continuous improvement of fleet provisioning services. * Participate in the Fleet Provisioning Automation on-call rotation and help drive incident response and post-incident improvements. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Empowering Thousands of Developers: Our Journey to an Internal Developer Platform](https://www.wearedevelopers.com/videos/1519-empowering-thousands-of-developers-our-journey-to-an-internal-developer-platform) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Answer the Interview Question: “Why Do You Want to Be a Software Engineer?”](https://www.wearedevelopers.com/magazine/392-how-to-answer-the-interview-question-why-do-you-want-to-be-a-software-engineer)