> Markdown version of [/jobs/ext/2072121-member-of-engineering-inference-infrastructure](https://www.wearedevelopers.com/jobs/ext/2072121-member-of-engineering-inference-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Member of Engineering (Inference Infrastructure) - **Company:** Poolside, Inc. - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Systems Engineering, Software Debugging, Distributed Systems, Fault Tolerance, Reinforcement Learning, Kubernetes - **Published:** August 15, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pb8ukzaymz ## About the Role * Strong programming skills in Go, or other similar languages * Strong systems engineering background: distributed systems, schedulers, control planes, or high-throughput data planes. * Production experience with Kubernetes internals - controllers, informers, operators - not just deploying to it. * Bias toward observability and debuggability: building a system that is easy to navigate when debugging production issues * Plus: experience in systems serving large scale inference requests ## Description You'll be working in the compute team focusing on GPU workload scheduling and inference serving optimization. You would partner with the inference team to improve our inference throughput and latency for evals and reinforcement learning. You would collaborate with our scalability team to focus on stabilizing our large scale fault tolerant training. You would also be in close contact with the infra team to make sure our GPU nodes are all healthy and fully utilized. We are one of the key teams to improve the research velocity. Any improvement on our systems has a wide impact on researchers and can contribute to the poolside mission on building a frontier model., * Design and develop internal scheduling system to maximize GPU utilization * Build API and tooling to help manage the lifecycle of GPU workloads and troubleshoot failures * Design and improve inference control plane to speed up model deployment and inference request serving * Collaborate with research to improve research velocity continuously ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [Using AI Without Losing Your Skills](https://www.wearedevelopers.com/videos/2045-using-ai-without-losing-your-skills) - [Throwing off the burdens of scale in engineering](https://www.wearedevelopers.com/videos/608-throwing-off-the-burdens-of-scale-in-engineering) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)