> Markdown version of [/jobs/ext/2000074-sr-software-engineer-ai-ml-aws-neuron-distributed-training](https://www.wearedevelopers.com/jobs/ext/2000074-sr-software-engineer-ai-ml-aws-neuron-distributed-training). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - **Company:** Amazon.com, Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Code Review, Computer Programming, Distributed Computing Environment, Machine Learning, Systems Development Life Cycle, Software Engineering, Pytorch, Large Language Models, Information Technology, HuggingFace, Machine Learning Operations - **Published:** August 9, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/sr-software-engineer-ai-ml-aws-neuron-distributed-training-cupertino-ca-usa-58872849 ## About the Role for stability across reduced precision * Profile, analyze, and tune end-to-end training pipelines for Trainium hardware * Collaborate with hardware, compiler, and runtime teams to influence system design and unlock new capabilities * Deploy and optimize training workloads at scale with AWS solution architects and customers Tasks * Bachelor's degree in computer science or equivalent * 5+ years of professional software development experience * 5+ years of programming experience in at least one language * 5+ years of design or architecture experience on large systems * 5+ years of full SDLC experience including code reviews, build, testing, and operations * Experience as a mentor, tech lead, or leading an engineering team * Experience in machine learning, large-scale training with LLMs and expertise in PyTorch Key requirements * health insurance * 401(k) matching * paid time off * RSUs * sign-on payments * parental leave ## Description Experteer Overview In this role, you will design and optimize distributed training for large-scale ML models on Trainium. You will work at the intersection of ML research and high-performance systems, collaborating with hardware, compiler, and runtime teams to deliver cost-effective, performant training on AWS Trainium. You will advance mixed-precision training and precision-aware strategies to maximize throughput while preserving accuracy. This is a hands-on opportunity to shape scalable ML pipelines and influence system design in a cutting-edge accelerator ecosystem. Compensation / Benefits * Design, implement and optimize distributed training solutions for large-scale ML models on Trainium instances * Extend and optimize distributed training frameworks (FSDP, torchtitan, Hugging Face) for the Neuron ecosystem * Develop and optimize mixed-precision and low-precision training techniques (BF16, FP8) * Implement precision-aware training strategies, loss scaling, and gradient management for stability across reduced precision * Profile, analyze, and tune end-to-end training pipelines for Trainium hardware * Collaborate with hardware, compiler, and runtime teams to influence system design and unlock new capabilities * Deploy and optimize training workloads at scale with AWS solution architects and customers Tasks * Bachelor's degree in computer science or equivalent * 5+ years of professional software development experience * 5+ years of programming experience in at least one language * 5+ years of design or architecture experience on large systems * 5+ years of full SDLC experience including code reviews, build, testing, and operations * Experience as a mentor, tech lead, or leading an engineering team * Experience in machine learning, large-scale training with LLMs and expertise in PyTorch Key requirements * health insurance * 401(k) matching * paid time off * RSUs * sign-on payments * parental leave ## Related Videos - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)