> Markdown version of [/jobs/ext/2712759-software-engineer-ai-inference-runtime-platform](https://www.wearedevelopers.com/jobs/ext/2712759-software-engineer-ai-inference-runtime-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer (AI Inference & Runtime Platform - **Company:** AZX INCORPORATED - **Location:** Seattle, WA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Software Debugging, Distributed Systems, Python (Programming Language), Key Management, Neo4j, Open Source Technology, Role-Based Access Control, Datadog, Autoscaling, Large Language Models, Backend, Fastapi, Kubernetes, Bicep, Terraform, Serverless Computing - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-ai-inference-runtime-platform-azx-9809439 ## About the Role * 5+ years of shipping production systems in a systems language. Rust is the house language, but polyglots are welcome - deep Go, C/C++, or Zig with genuine appetite for Rust counts. Async runtimes, memory-safety discipline, and debugging at the syscall boundary should be familiar territory. * Operated Kubernetes workloads that other people depended on - controllers or operators, scheduling, autoscaling, node lifecycle. You've been paged, and the experience changed how you build. * Strong ability to threat-model isolation boundaries (namespaces, cgroups, seccomp, hypervisors), including identifying what an untrusted guest could observe, forge, or exhaust, and applying security best practices for agentic execution - least privilege, no credentials in the sandbox, audit trails, and human approval on write actions. * Hands-on experience deploying or operating open-weight LLM serving infrastructure (vLLM/SGLang or similar), including packaging models into reliable, metered production endpoints. * Performance discipline in distributed systems: you measure before you optimize, and you can tell the story of a latency you killed with the numbers attached. * Practical depth in some of our core stack - Rust (tokio), Python/FastAPI, Kubernetes operators (controller-runtime/Kubebuilder/CRDs), KEDA, Karpenter, GPU device plugins/DRA - with a genuine willingness to research your way into the rest. * Familiarity with isolation technology (Firecracker, Kata, gVisor, or comparable), secrets management and egress control (Vault/KMS-class), and hosting stateful systems (vector stores like pgvector/Qdrant, graph stores like Neo4j) with backup and failover discipline. * Comfort operating across cloud and GPU substrates - AWS/Azure/GCP plus managed GPU clouds - using infrastructure-as-code (Terraform/OpenTofu, Bicep) and observability tooling (OpenTelemetry). * Bachelor's Degree; Master's is a plus ## Description You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. This role anchors on two systems. The first is our inference control plane - open-weight models and custom task-model zoos, hosted and operated across managed GPU clouds and customer-managed Kubernetes clusters, with scale-to-zero economics, cold-start discipline, and per-token cost accounting that stays correct even when a client disconnects mid-stream - along with the Kubernetes layer those workloads live on: operators, autoscaling, node lifecycle. The second is our agent-sandboxing platform: hardware-isolated microVMs for running untrusted, agent-generated code securely and compliantly by construction, where agents operate with least privilege, never see a credential, and a human gates anything that writes to a system of record. You'll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home - building the fork engine, the guest agent, and the multi-substrate model lifecycle. We're looking for individuals who've built this class of stack (an inference-serving or serverless-GPU platform), operated it hard at scale, or ideally both. Responsibilities: * Manage the serving tier for open-weight models: engine deployment and configuration, cold-start strategy, per-model SLOs, and upgrade/canary discipline. * Administer the Kubernetes layer for inference and sandbox workloads: operators and CRDs, autoscaling (KEDA/Karpenter-class), GPU scheduling and sharing, and node lifecycle. * Own the stateful data plane: hosted vector stores for semantic memory and graph stores for knowledge graphs - deployed, backed up, scaled, and recovered, with restores that are tested rather than hoped for. * Oversee the sandbox runtime and its host-side control plane: lifecycle, exec, snapshot/fork, teardown, metering, and the threat model of the isolation boundary. * Direct the FastAPI control-plane services, Terraform/OpenTofu, Bicep, and the dashboards * Run the layer the backend services team builds on, expect to debug into their services, and expect them to read your dashboards. * Manage the open-source posture: build to OSS standards and release as it matures, with reviewed PRs, real docs, and reproducible builds. ## Related Videos - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Building AI Solutions with Rust and Docker](https://www.wearedevelopers.com/magazine/494-building-ai-solutions-with-rust-and-docker) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)