> Markdown version of [/jobs/ext/1921976-senior-ml-software-engineer-data-plane](https://www.wearedevelopers.com/jobs/ext/1921976-senior-ml-software-engineer-data-plane). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Software Engineer, Data Plane - **Company:** Amazon.com, Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Salary:** $193,300.0 - $261,500.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Code Review, Extract Transform Load (ETL), Linux, Distributed Systems, Memory Management, Machine Learning, Open Source Technology, Remote Direct Memory Access, Tensorflow, Software Engineering, Network Switches, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Model Validation, Information Technology, Optimization Algorithms, Build Process, TensorRT, Software Coding, Decoding, Software Version Control - **Published:** August 4, 2026 - **Apply:** https://www.amazon.jobs/en/jobs/10491192/senior-ml-software-engineer-data-plane ## About the Role Bachelor's degree in computer science or equivalent - 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Strong proficiency in C/C++ - Strong Linux systems knowledge - Experience developing compute kernels for GPUs, DSPs, or custom accelerators - Proven track record of owning and delivering complex software features end-to-end, Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT - Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations - Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming - Experience with hardware simulation environments and model validation workflows - Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow ## Description The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration. Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models. This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack. Key job responsibilities - Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference. - Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware. - Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism. - Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets. - Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup. - Own features end-to-end: from design through implementation, testing, and integration into the broader software stack. - Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions. - Mentor engineers, drive design reviews, and raise the engineering bar across the team. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)