> Markdown version of [/jobs/ext/2846297-staff-software-engineer](https://www.wearedevelopers.com/jobs/ext/2846297-staff-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - **Company:** Coreweave, Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $188,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Program Optimization, Computer Programming, Distributed Systems, Python (Programming Language), Performance Tuning, Service Design, Large Language Models, Caching, Reliability of Systems, Kubernetes, Low Latency, TensorRT - **Published:** September 11, 2026 - **Apply:** https://www.thejobnetwork.com/job/78a2ffa5-f955-4a55-bef5-030e3846eb23/staff-software-engineer-inference ## About the Role - 8-12+ years of experience building and operating large-scale distributed systems or cloud platforms - Proven experience leading cross-team technical initiatives impacting multiple services or organizations - Strong programming skills in Go, Python, or C++ - Deep expertise in Kubernetes at production scale, including orchestration, scheduling, and service design - Strong understanding of distributed systems, networking, and performance optimization - Experience designing and operating low-latency, high-throughput systems with strict P95/P99 latency requirements - Hands-on experience with inference systems, including batching or micro-batching strategies, caching, and memory optimization - Experience improving system performance using metrics-driven approaches (e.g., latency, throughput, utilization) - Familiarity with mixed precision (BF16, FP8) and streaming inference workloads **Preferred:** - Experience with inference frameworks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe ## Description **About the role:** As a Staff Software Engineer (IC5) on the Inference team, you will act as a technical leader driving architecture, performance, and reliability across multiple services and teams. Your day-to-day will involve leading cross-team design initiatives, optimizing inference performance (latency, throughput, and GPU utilization), and improving system reliability at scale. You will work deeply in distributed systems and Kubernetes-based infrastructure, focusing on areas like scheduling, batching, and memory optimization. This role requires hands-on technical leadership and the ability to influence engineering direction across the organization. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)