> Markdown version of [/jobs/ext/2714958-ai-inference-infrastructure-software-engineer](https://www.wearedevelopers.com/jobs/ext/2714958-ai-inference-infrastructure-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Infrastructure Software Engineer - **Company:** Elastixai Inc. - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Bash Shell, System Configuration, Linux, Field-Programmable Gate Array (FPGA), Identity and Access Management, Python (Programming Language), Octopus Deploy, Reliability Engineering, Ansible, Software Engineering, Rust (Programming Language), Pulumi, Graphics Processing Unit (GPU), Google Cloud, Autoscaling, Kubernetes, Information Technology, Hardware Infrastructure, Puppet, Software Coding, Terraform, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/ai-inference-infrastructure-software-engineer-kubernetes-cloud-elastixai-inc-8300788 ## About the Role * Minimum BS in Computer Science, Software Engineering, or a related field. * 3-5 years of hands-on Kubernetes experience, including EKS, GKE, and/or self-hosted clusters. * 2-3 years of production experience operating workloads on AWS or GCP. * Proven track record running ML or inference services at scale on Kubernetes in production. * Strong experience running accelerated workloads in Kubernetes, scheduling, drivers, device plugins, MIG, networking, and storage considerations. * Solid coding skills in Python, Bash and proficiency in Go * Proficient in configuring and leveraging Linux OS in production * Experience with infrastructure-as-code (Terraform, Pulumi), OS configuration state (Ansible, Puppet, Salt) and GitOps workflows (Argo CD, Flux). * Experience in OS configuration tooling. * Familiarity with AI inference and/or training workflows and the operational patterns around them. * Pragmatic, ownership-oriented mindset; comfortable operating in early-stage ambiguity and shipping iteratively. Preferred/Bonus Qualifications: * MS/PhD in Computer Science, Software Engineering, or a related field. * Experience with inference servers and runtimes (e.g., Triton, vLLM, TGI) and model-serving patterns (batching, streaming, KV-cache aware routing). * Exposure to heterogeneous accelerators beyond GPUs (FPGAs, custom ASICs). * Background in observability, SRE, or performance engineering for latency-sensitive services. * Experience building customer facing API platforms including onboarding, API keys/auth management, and usage metering. ## Description This is a hands-on role with broad surface area. You'll touch everything from cluster bring-up, automating the software releases, and AI Accelerator scheduling to service reliability and cost optimization, working closely with our ML, runtime, and hardware teams to expose the full performance of our co-designed stack to end users., * Build, operate, and evolve ElastixAI's Kubernetes infrastructure powering our Token-as-a-Service capability. * Run accelerated inference workloads in production at scale, with strong SLAs around latency, throughput, and availability. * Manage and harden our AWS, GCP, and on-prem infrastructure, including networking, storage, IAM, and observability layers tied to our services. * Develop tooling and automation in Python, Bash, Rust, and Go to streamline deployments, rollouts, autoscaling, and incident response. * Partner with the ML and runtime teams to productionize new inference capabilities, model deployments, and routing strategies. * Contribute to capacity planning, cost optimization, and reliability engineering across multi-cloud and self-hosted environments. * Help define the platform roadmap as we scale from early customers to broad production deployments. * Be a member of the Elastix On-Call rotation ## Related Videos - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)