> Markdown version of [/jobs/ext/2726665-ml-systems-engineer-inference-acceleration](https://www.wearedevelopers.com/jobs/ext/2726665-ml-systems-engineer-inference-acceleration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Systems Engineer - Inference Acceleration - **Company:** Arago - **Location:** Courbevoie, France - **Contract:** Permanent contract - **Skills:** Computer-Aided Design, C++ (Programming Language), Nvidia CUDA, Computer Programming, Extract Transform Load (ETL), Python (Programming Language), System Programming, Large Language Models, Caching, Parallel Computation, Machine Learning Operations, TensorRT - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/ml-systems-engineer-inference-acceleration-arago-9013268 ## About the Role * Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering. * Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks. * Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments. * Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads. * Strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/computation overlap. * Hands-on experience with modern inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling. * Strong C++ and Python skills, and comfort working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively being developed. Exposure to or experience with MLIR and MLIR dialects is a strong plus. * Language: English at a proficient level. ## Description * Analyze modern AI workloads and identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator. * Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization. * Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies. * Develop inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation. * Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads. * Work closely with Arago's hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features based on real model workloads. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market)