> Markdown version of [/jobs/ext/2709087-software-engineer-ai-inference-platform](https://www.wearedevelopers.com/jobs/ext/2709087-software-engineer-ai-inference-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, AI Inference Platform - **Company:** Elastixai Inc. - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Abstraction Layers, Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Computer Engineering, Shard (Database Architecture), Software Debugging, Distributed Systems, Memory Management, Design of User Interfaces, Python (Programming Language), Open Source Technology, PCI Express, Performance Tuning, System Programming, Pytorch, Large Language Models, Containerization, Kubernetes, Information Technology, Low Latency, Machine Learning Operations, TensorRT, Docker - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/ai-software-engineer-elastixai-inc-8300791 ## About the Role * BS/MS/PhD in Computer Science, Electrical/Computer Engineering, or related field. * 3+ years of professional experience in systems programming, ML infrastructure, or distributed inference. * Proficient in C++ and Python, with strong debugging and performance analysis skills. * Deep familiarity with one or more LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, DeepSpeed-Inference, etc.). * Understanding of model deployment internals - token scheduling, KV caching, batching, and pipelined inference. * Comfortable working close to the hardware abstraction layer - CUDA, PCIe, memory management, or runtime scheduling. * Strong collaboration and communication skills; ability to work cross-functionally in a fast-paced startup environment., * Experience with hardware-aware ML optimization, compiler/runtime integration, or accelerator SDKs. * Hands-on experience profiling GPU/accelerator workloads. * Familiarity with containerized deployments (Docker/Kubernetes). * Exposure to distributed systems or large-scale inference clusters. * Contributions to open-source ML or serving frameworks. ## Description We're looking for a systems-minded AI Software Engineer to join our core inference platform team. You'll design and extend the low-level serving stack - hacking open-source frameworks like vLLM, SGLang, and TensorRT-LLM, building new model sharding and scheduling logic, and integrating deeply with our proprietary AI accelerator. This role sits at the intersection of ML systems, compiler/runtime engineering, and hardware-software co-design., * Architect, extend, and optimize core components of our AI serving platform for throughput, latency, and scalability. * Customize open-source serving frameworks (e.g., vLLM) for proprietary model ingestion and accelerator integration. * Develop efficient model partitioning, scheduling, and memory management strategies for multi-device inference. * Collaborate with ML engineers on model export and runtime optimization (quantization, graph transforms). * Work closely with hardware engineers to influence accelerator interface design and performance tuning. * Build APIs and runtime tools enabling flexible, PyTorch-native model deployment on our infrastructure. * Profile, debug, and optimize across the full stack - from Python orchestration to C++ kernels and PCIe drivers. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)