> Markdown version of [/jobs/ext/2727528-artificial-general-intelligence](https://www.wearedevelopers.com/jobs/ext/2727528-artificial-general-intelligence). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Artificial General Intelligence - **Company:** Agi - **Location:** UK (Remote available) - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Systems Engineering, Software Debugging, Distributed Systems, Fault Tolerance, Reinforcement Learning, Kubernetes - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/member-of-engineering-inference-infrastructure-poolside-8752108 ## About the Role * Strong programming skills in Go, or other similar languages * Strong systems engineering background: distributed systems, schedulers, control planes, or high-throughput data planes. * Production experience with Kubernetes internals - controllers, informers, operators - not just deploying to it. * Bias toward observability and debuggability: building a system that is easy to navigate when debugging production issues * Plus: experience in systems serving large scale inference requests ## Description You'll be working in the compute team focusing on GPU workload scheduling and inference serving optimization. You would partner with the inference team to improve our inference throughput and latency for evals and reinforcement learning. You would collaborate with our scalability team to focus on stabilizing our large scale fault tolerant training. You would also be in close contact with the infra team to make sure our GPU nodes are all healthy and fully utilized. We are one of the key teams to improve the research velocity. Any improvement on our systems has a wide impact on researchers and can contribute to the poolside mission on building a frontier model., * Design and develop internal scheduling system to maximize GPU utilization * Build API and tooling to help manage the lifecycle of GPU workloads and troubleshoot failures * Design and improve inference control plane to speed up model deployment and inference request serving * Collaborate with research to improve research velocity continuously ## Related Videos - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [Using AI Without Losing Your Skills](https://www.wearedevelopers.com/videos/2045-using-ai-without-losing-your-skills) - [Agentic AI in Go](https://www.wearedevelopers.com/videos/100274-agentic-ai-in-go) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Bringing AI Everywhere](https://www.wearedevelopers.com/videos/1132-bringing-ai-everywhere) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)