> Markdown version of [/jobs/ext/2872607-software-development-engineer](https://www.wearedevelopers.com/jobs/ext/2872607-software-development-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Development Engineer - **Company:** Amazon.com, Inc. - **Location:** Monte Vista, CA, United States - **Experience:** Experienced - **Salary:** $165,200.0 - $223,600.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Code Review, Nvidia CUDA, Software Debugging, Software Design Patterns, Memory Management, Python (Programming Language), Linux Kernel, Machine Learning, Software Engineering, Pytorch, Large Language Models, Parallel Computation, Information Technology, TensorRT - **Published:** September 13, 2026 - **Apply:** https://www.thejobnetwork.com/job/cd2119ab-8a41-4ab3-938a-6d794fd39524/software-development-engineer-aiml-aws-neuron-model-inference ## About the Role nThe Inference Enablement and Acceleration team fosters a builder's culture where experimentation is encouraged, and impact is measurable. We emphasize collaboration, technical ownership, and continuous learning. Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future. Join us to solve some of the most interesting and impactful infrastructure challenges in AI/ML today. \n \nBASIC QUALIFICATIONS- Bachelor's degree in computer science or equivalent\n \n- 3+ years of non-internship professional software development experience\n \n- 3+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience\n \n- Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model execution.\n \n- Software development experience in C++, Python (experience in at least one language is required).\n \n- Strong understanding of system performance, memory management, and parallel computing principles.\n \n- Proficiency in debugging, profiling, and implementing best software engineering practices in large-scale systems.\n \nPREFERRED QUALIFICATIONS- Familiarity with PyTorch, JIT compilation, and AOT tracing.\n \n- Familiarity with CUDA kernels or equivalent ML or low-level kernels\n \n- Candidates with performant kernel development such as CUTLASS, FlashInfer etc., would be well suited.\n \n- Familiar with syntax and tile-level semantics similar to Triton.\n \n- Experience with online/offline inference serving with vLLM, SGLang, TensorRT or similar platforms in production environments.\n \n- Deep understanding of computer architecture, operation systems level software and working knowledge of parallel computing.\n ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)