> Markdown version of [/jobs/ext/2704209-software-engineer-ai-inference-platform](https://www.wearedevelopers.com/jobs/ext/2704209-software-engineer-ai-inference-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, AI Inference Platform - **Company:** Elastixai Inc. - **Location:** Seattle, WA, United States - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Code Generation, Data Structures, Field-Programmable Gate Array (FPGA), Python (Programming Language), Machine Learning, Open Source Technology, Tensorflow, Software Construction, Software Engineering, Systems Architecture, Application Specific Integrated Circuits, Pytorch, Large Language Models, Deep Learning, Information Technology, Hardware Acceleration, Machine Learning Operations, Programming Languages - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/ai-compiler-and-performance-engineer-elastixai-inc-8300786 ## About the Role * BS/MS/PhD in Computer Science, Software Engineering, or a related field. * Deep experience building compilers, optimizing kernels, or working with ML frameworks at a systems level. * Strong proficiency in one or more programming languages such as Python and C++. * Strong understanding of one or more of the following: * LLM architectures and transformer internals * MLIR, LLVM, XLA, TVM, Triton, or similar compiler infrastructures * GPU/TPU/FPGA/ASIC compute models, memory hierarchies, and parallel execution * Quantization, sparsity, or algorithmic optimization for deep learning * Deep expertise on ML frameworks (e.g., PyTorch, TensorFlow, JAX) and understanding of ML model deployment challenges. * Solid understanding of software engineering best practices, including data structures, algorithms, and testing. * Thinking in terms of latency, cycles, memory bandwidth, and arithmetic intensity, not just algorithms. * Excellent problem-solving abilities and a knack for tackling complex technical challenges. * Excited to collaborate across ML, hardware, and software boundaries to invent something fundamentally new. * Strong communication skills and a proven ability to collaborate effectively in a cross-functional team environment. * Ability to thrive in a fast-paced, dynamic startup environment., * PhD in Computer Science, Software Engineering, or a related field. * Experience with custom hardware accelerators for ML inference. * Contributions to open-source compiler or ML systems projects. * Prior startup experience or background building first-generation systems. ## Description * Break down LLM and transformer workloads into fine-grained primitives tailored to our proprietary compute hardware. * Design and implement IR transformations, graph optimizations, kernel lowering, and code generation for novel hardware architectures. * Collaborate with ML researchers to co-design algorithmic optimizations that yield real end-to-end performance gains. * Work closely with hardware architects to refine microarchitectural features, instruction sets, memory hierarchies, and execution models. * Build performance models, profiling tools, and benchmarking frameworks to identify bottlenecks and guide design decisions. * Prototype and validate improvements across the entire stack - from PyTorch/XLA-level passes to custom kernel implementations. * Contribute to shaping the overall system architecture of a first-of-its-kind inference engine. ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)