> Markdown version of [/jobs/ext/1401070-sr-software-engineer-ai-ml-aws-neuron-distributed-training](https://www.wearedevelopers.com/jobs/ext/1401070-sr-software-engineer-ai-ml-aws-neuron-distributed-training). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - **Company:** Amazon.com, Inc. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $168,100.0 - $227,400.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, Amazon Web Services, Code Review, Computer Programming, Software Design Patterns, Distributed Computing Environment, Machine Learning, Tensorflow, Software Engineering, Pytorch, Large Language Models, Information Technology, Build Process, Software Coding, Software Version Control - **Published:** July 23, 2026 - **Apply:** https://dejobs.org/x/x/EE626B3D6EC446DC91775B7CF8340621/job/ ## About the Role * Bachelor's degree in computer science or equivalent * 5+ years of non-internship professional software development experience * 5+ years of programming with at least one software programming language experience * 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience * 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience * Experience as a mentor, tech lead or leading an engineering team * Experience in machine learning, large scale training with LLMs and expertise in Pytorch., * Master's degree in computer science or equivalent * Experience in computer architecture * Previous software engineering expertise with Pytorch/Jax/Tensorflow, Distributed libraries and Frameworks, End-to-end Model Training. ## Description The ML Distributed Training team works side by side with chip architects, compiler engineers and runtime engineers to create, build and tune distributed training solutions with Trainium instances. Experience with training these large models using Pythorch is a must. Distributed training with awareness of strategies like FSDP (Fully-Sharded Data Parallel), PP, Context parallel. Distributed training libraries like torchtitan, torchtune , HF RL , DeepSeek etc are central to this and extending all of this for the Neuron based system is key focussing on enabling large scale training. Experience is post-training strategies like DPO/PPO/HF torch-tune will additional strength and aligns with team success., You will lead efforts to build distributed training support into PyTorch, the Neuron compiler, and runtime stacks. You will enable distribute training strategies as well as use them to optimize models to achieve peak performance and maximize efficiency on AWS custom silicon, including Trainium servers. Strong software development skills, the ability to deep dive, work effectively within cross-functional teams, and a solid foundation in Machine Learning are critical for success in this role. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)