> Markdown version of [/jobs/ext/1971244-ml-software-engineer-data-plane](https://www.wearedevelopers.com/jobs/ext/1971244-ml-software-engineer-data-plane). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Software Engineer, Data Plane - **Company:** Amazon.com, Inc. - **Location:** Springfield, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Computer Literacy, Extract Transform Load (ETL), Linux, Memory Management, Open Source Technology, Software Engineering, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Parallel Computation - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/ml-software-engineer-data-plane-illinois-usa-58834390 ## About the Role Experteer Overview In this role you will own and optimize the inference data plane for a custom ML accelerator, enabling efficient large-model execution across hardware, memory, and data movement. You'll work closely with cross-functional teams to bring models from validation to production, shaping high-performance kernels and integration with serving frameworks. You'll scale and validate end-to-end architectures, building test and profiling tools to drive latency and throughput. This ground-up role offers impact across the full stack and opportunities to influence frontier-scale ML deployments. Compensation / Benefits * Develop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performance * Implement and validate end-to-end LLM architectures from PyTorch models to distributed execution on custom hardware * Integrate custom accelerator backends into open-source serving frameworks (vLLM, PyTorch) including scheduler extensions, memory aaa backends and model parallelism * Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets * Profile and optimize inference workloads, identify bottlenecks, and drive latency/throughput improvements from simulation to hardware bring-up * Own features end-to-end from design through implementation, testing, and integration into the software stack * Contribute to CI/CD pipelines to gate model and kernel changes on correctness and performance regressions Tasks * Bachelor's degree or equivalent * 4+ years of full software development lifecycle experience * Knowledge of computer architecture, operating systems, and parallel computing * Strong proficiency in C/C++ * Strong Linux systems knowledge * Experience developing compute kernels for GPUs, DSPs, or custom accelerators * Proven track record of owning and delivering complex software features end-to-end Key requirements * ## Description Experteer Overview In this role you will own and optimize the inference data plane for a custom ML accelerator, enabling efficient large-model execution across hardware, memory, and data movement. You'll work closely with cross-functional teams to bring models from validation to production, shaping high-performance kernels and integration with serving frameworks. You'll scale and validate end-to-end architectures, building test and profiling tools to drive latency and throughput. This ground-up role offers impact across the full stack and opportunities to influence frontier-scale ML deployments. Compensation / Benefits * Develop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performance * Implement and validate end-to-end LLM architectures from PyTorch models to distributed execution on custom hardware * Integrate custom accelerator backends into open-source serving frameworks (vLLM, PyTorch) including scheduler extensions, memory management, and model parallelism * Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets * Profile and optimize inference workloads, identify bottlenecks, and drive latency/throughput improvements from simulation to hardware bring-up * Own features end-to-end from design through implementation, testing, and integration into the software stack * Contribute to CI/CD pipelines to gate model and kernel changes on correctness and performance regressions Tasks * Bachelor's degree or equivalent * 4+ years of full software development lifecycle experience * Knowledge of computer architecture, operating systems, and parallel computing * Strong proficiency in C/C++ * Strong Linux systems knowledge * Experience developing compute kernels for GPUs, DSPs, or custom accelerators * Proven track record of owning and delivering complex software features end-to-end Key requirements * ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Tackling the Risks of AI - With AI](https://www.wearedevelopers.com/videos/1690-tackling-the-risks-of-ai-with-ai) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)