Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
Experteer Overview In this role, you will design and optimize distributed training for large-scale ML models on Trainium. You will work at the intersection of ML research and high-performance systems, collaborating with hardware, compiler, and runtime teams to deliver cost-effective, performant training on AWS Trainium. You will advance mixed-precision training and precision-aware strategies to maximize throughput while preserving accuracy. This is a hands-on opportunity to shape scalable ML pipelines and influence system design in a cutting-edge accelerator ecosystem. Compensation / Benefits * Design, implement and optimize distributed training solutions for large-scale ML models on Trainium instances * Extend and optimize distributed training frameworks (FSDP, torchtitan, Hugging Face) for the Neuron ecosystem * Develop and optimize mixed-precision and low-precision training techniques (BF16, FP8) * Implement precision-aware training strategies, loss scaling, and gradient management for stability across reduced precision * Profile, analyze, and tune end-to-end training pipelines for Trainium hardware * Collaborate with hardware, compiler, and runtime teams to influence system design and unlock new capabilities * Deploy and optimize training workloads at scale with AWS solution architects and customers Tasks * Bachelor’s degree in computer science or equivalent * 5+ years of professional software development experience * 5+ years of programming experience in at least one language * 5+ years of design or architecture experience on large systems * 5+ years of full SDLC experience including code reviews, build, testing, and operations * Experience as a mentor, tech lead, or leading an engineering team * Experience in machine learning, large-scale training with LLMs and expertise in PyTorch Key requirements * health insurance * 401(k) matching * paid time off * RSUs * sign-on payments * parental leave
Requirements
for stability across reduced precision * Profile, analyze, and tune end-to-end training pipelines for Trainium hardware * Collaborate with hardware, compiler, and runtime teams to influence system design and unlock new capabilities * Deploy and optimize training workloads at scale with AWS solution architects and customers Tasks * Bachelor’s degree in computer science or equivalent * 5+ years of professional software development experience * 5+ years of programming experience in at least one language * 5+ years of design or architecture experience on large systems * 5+ years of full SDLC experience including code reviews, build, testing, and operations * Experience as a mentor, tech lead, or leading an engineering team * Experience in machine learning, large-scale training with LLMs and expertise in PyTorch Key requirements * health insurance * 401(k) matching * paid time off * RSUs * sign-on payments * parental leave
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are Large Language Models?
MLOps And AI Driven Development
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence