> Markdown version of [/jobs/ext/2846690-specialist-field-engineer-compute-infrastructure](https://www.wearedevelopers.com/jobs/ext/2846690-specialist-field-engineer-compute-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Specialist Field Engineer - Compute Infrastructure - **Company:** Coreweave Inc - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $188,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Bash Shell, BIOS, Command-Line Interface, Cloud Computing, Computer Clusters, Data Centers, Distributed Systems, Firmware, InfiniBand, Python (Programming Language), Linux System Administration, Routing, Node.Js, Raw Data, Cloud Services, Ansible, Supercomputing, TCP/IP, Graphics Processing Unit (GPU), Performance Testing, Multi-Cloud, Kubernetes, Information Technology, Bare Metal, Slurm, Hardware Infrastructure - **Published:** September 11, 2026 - **Apply:** https://www.thejobnetwork.com/job/c7572c67-e473-49d1-87ae-5d9f21815744/senior-specialist-field-engineer-compute-infrastructure ## About the Role * B.S. in Computer Science or a related technical discipline, or equivalent experience * 7+ years of proven experience as a Solutions Architect, Field Engineer, Infrastructure/Systems Engineer, or Technical Account Manager in Cloud Infrastructure, focusing on building or operating distributed systems or HPC/cloud services, with an expertise focused on bare-metal compute infrastructure and large-scale GPU cluster delivery * Fluency in cloud computing concepts, architecture, and technologies with hands-on experience in designing and implementing cloud solutions * Proven track record with building customer relationships, communicating clearly and the ability to break down complex technical concepts to both technical and non-technical audiences * Deep expertise with modern rack-scale GPU server hardware (e.g., NVIDIA HGX / GB200-class systems), high-speed interconnects (InfiniBand, NVLink), and the firmware/BMC/BIOS layer * Expert-level Linux system administration and command-line troubleshooting, paired with strong networking fundamentals (routing, fabric topologies, TCP/IP) * Hands-on experience bringing up, validating, and operating large GPU clusters-including bare metal node pxe boot, hardware health, fabric validation, and HPC acceptance/performance testing-and integrating bare metal with orchestration layers such as Kubernetes and Slurm Preferred: * Experience operating security-sensitive, air-gapped, or otherwise locked-down customer environments * Experience with scripting and automation related to bare-metal provisioning, infrastructure validation, and lifecycle management (Python, Bash, Ansible, or similar) * Experience designing AI supercomputers from MEP designs * Experience delivering bare-metal infrastructure at scale for large strategic customers or AI research labs * Experience with building solutions across multi-cloud or hybrid environment Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You love to help solve challenging technical problems * You're curious about the latest and greatest technologies in the AI space * You're an expert in managing conflict and achieving mutually beneficial technical outcomes ## Description As a Specialist Field Engineer - Compute Infrastructure at CoreWeave, you'll own the technical path for some of our largest customers as they go from facility and rack design to a validated, production-ready supercomputer. Working alongside the teams that build and operate each layer, you are the deep technical expert who turns raw data center hardware-racks, GPUs, high-speed fabric, firmware-into reliable compute that customers can train and inference on at scale, spanning infrastructure engineering, provisioning, validation, operations, and support. You'll engage hands-on across the entire customer lifecycle: leading new GPU cluster bring-up and acceptance, driving InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we operate customer bare-metal fleets at rack-level-and-up (IT service, break-fix, network, and firmware), and standing up locked-down, security-sensitive environments for our most strategic AI customers. You'll partner closely with Data Center Operations, Fleet Operations, Networking, and Product Engineering, and your work in the field will directly shape how CoreWeave delivers compute infrastructure. If you're driven by innovation, thrilled by the possibilities of what specialized compute can enable, and eager to be part of a team that's shaping the future, then CoreWeave is the place for you. Join us and let's embark on this adventure together! In this role, you will: * Serve as the primary technical point of contact for customers, establishing strong technical relationships and ensuring their success with CoreWeave's cloud infrastructure offerings, focusing on bare-metal compute infrastructure and end-to-end cluster delivery within high-performance compute (HPC) environments. * Own the technical path from facility and rack design to a validated, production-ready supercomputer-spanning logical design, infrastructure engineering, provisioning, validation, operations, and support. * Lead bring-up and acceptance of new large-scale GPU clusters, driving InfiniBand/RoCE fabric validation, HPC performance benchmarking (e.g., NCCL, ib_write_bw), and remediation of fabric, optics, firmware, and node-level issues to meet customer performance targets. * Define and operationalize models for managing customer bare-metal fleets at rack-level-and-up-IT service, break-fix, network and firmware management-including Bare Metal as a Service (BMaaS) and customer self-service patterns. * Partner with Data Center Operations, Fleet Operations, and Networking teams to align facility, hardware, and fabric readiness with customer go-live timelines and operational SLAs. * Review and advise on customer-facing technical contract terms, including service scope, operational responsibilities, SLAs, isolation requirements, and support boundaries. * Lead proof of concept initiatives to showcase the value and viability of CoreWeave's solutions within specific environments. * Drive technical leadership and direction during customer meetings, presentations, and workshops, addressing any technical queries or concerns that arise. * Act as a virtual member of CoreWeave's Compute Infrastructure, Fleet Operations, and Networking engineering teams, identifying opportunities for product enhancement and collaborating with engineers to implement your suggestions. * Offer valuable insights on product features, functionality, and performance, contributing regularly to discussions about product strategy and architecture. * Stay informed of the latest developments and trends in Kubernetes, cloud computing and infrastructure, sharing your thought leadership with customers and internal stakeholders. * Lead the prototyping and initiation of research and development efforts for emerging products and solutions, delivering prototypes and key insights for internal consumption. * Represent CoreWeave at conferences and industry events, with occasional travel as required. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025)