Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
Experteer Overview In this role, you will design and optimize distributed training for large-scale ML models on Trainium. You will work at the intersection of ML research and high-performance systems, collaborating with hardware, compiler, and runtime teams to deliver cost-effective, performant training on AWS Trainium. You will advance mixed-precision training and precision-aware strategies to maximize throughput while preserving accuracy. This is a hands-on opportunity to shape scalable ML pipelines and influence system design in a cutting-edge accelerator ecosystem. Compensation / Benefits * Design, implement and optimize distributed training solutions for large-scale ML models on Trainium instances * Extend and optimize distributed training frameworks (FSDP, torchtitan, Hugging Face) for the Neuron ecosystem * Develop and optimize mixed-precision and low-precision training techniques (BF16, FP8) * Implement precision-aware training strategies, loss scaling, and gradient management for stability across reduced precision * Profile, analyze, and tune end-to-end training pipelines for Trainium hardware * Collaborate with hardware, compiler, and runtime teams to influence system design and unlock new capabilities * Deploy and optimize training workloads at scale with AWS solution architects and customers Tasks * Bachelor’s degree in computer science or equivalent * 5+ years of professional software development experience * 5+ years of programming experience in at least one language * 5+ years of design or architecture experience on large systems * 5+ years of full SDLC experience including code reviews, build, testing, and operations * Experience as a mentor, tech lead, or leading an engineering team * Experience in machine learning, large-scale training with LLMs and expertise in PyTorch Key requirements * health insurance * 401(k) matching * paid time off * RSUs * sign-on payments * parental leave
Requirements
for stability across reduced precision * Profile, analyze, and tune end-to-end training pipelines for Trainium hardware * Collaborate with hardware, compiler, and runtime teams to influence system design and unlock new capabilities * Deploy and optimize training workloads at scale with AWS solution architects and customers Tasks * Bachelor’s degree in computer science or equivalent * 5+ years of professional software development experience * 5+ years of programming experience in at least one language * 5+ years of design or architecture experience on large systems * 5+ years of full SDLC experience including code reviews, build, testing, and operations * Experience as a mentor, tech lead, or leading an engineering team * Experience in machine learning, large-scale training with LLMs and expertise in PyTorch Key requirements * health insurance * 401(k) matching * paid time off * RSUs * sign-on payments * parental leave
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are Large Language Models?
MLOps And AI Driven Development
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence