> Markdown version of [/jobs/ext/3115154-software-engineer-compute-architecture](https://www.wearedevelopers.com/jobs/ext/3115154-software-engineer-compute-architecture). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Compute Architecture - **Company:** Coreweave Inc - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $153,000.0 - $242,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Build Automation, Data Centers, Distributed Systems, Firmware, IBM Hardware Management Console, Open Source Technology, Prometheus, Datadog, Graphics Processing Unit (GPU), Grafana, Backend, Containerization, Kubernetes, Information Technology, Restful APIs, Network Server - **Published:** September 27, 2026 - **Apply:** https://www.juju.com/job/16_b9c180854 ## About the Role * 5+ years of experience building and operating infrastructure or backend systems. * Bachelor's or Master's degree in Computer Science or a related field, or equivalent practical experience. * Strong proficiency in Go for building production services and tools. * Experience designing and building gRPC and REST APIs. * Experience with Kubernetes and containerized workloads in production environments. * Familiarity with observability tooling such as Prometheus and Grafana. Preferred * Experience working with GPU-based systems. * Experience with low-level hardware management such as BMCs or Redfish. * Experience operating large-scale distributed systems or high-throughput infrastructure. * Experience collaborating with or contributing to open-source projects (for example, Go, Redfish). Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You enjoy working close to the hardware and are curious about how GPUs, servers, and data centers fit together. * You thrive in infrastructure environments where reliability, performance, and automation matter as much as features. * You like collaborating across hardware, platform, and product teams to solve complex, ambiguous problems. ## Description As a Senior Software Engineer within our Compute Architecture organization, you will help build the software control plane for hardware lifecycle management across large-scale GPU data centers. The METALDEV team builds Go-based distributed services that bring infrastructure online, monitor production hardware health, automate safe operational workflows, and give operators the observability and control needed to manage GPU servers and rack-scale systems with reliability and confidence. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, ideal for engineers who want their code to operate real-world infrastructure at massive scale. What You'll Do * Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure. * Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations. * Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure. * Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly. * Translate incidents and hardware failure modes into software improvements that make the platform more resilient. * Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)