Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Role details
Job location
Tech stack
Job description
You will lead efforts to build distributed training support into PyTorch, the Neuron compiler, and runtime stacks. You will enable distribute training strategies as well as use them to optimize models to achieve peak performance and maximize efficiency on AWS custom silicon, including Trainium servers. Strong software development skills, the ability to deep dive, work effectively within cross-functional teams, and a solid foundation in Machine Learning are critical for success in this role.
Requirements
The ML Distributed Training team works side by side with chip architects, compiler engineers and runtime engineers to create, build and tune distributed training solutions with Trainium instances. Experience with training these large models using Pythorch is a must. Distributed training with awareness of strategies like FSDP (Fully-Sharded Data Parallel), PP, Context parallel. Distributed training libraries like torchtitan, torchtune, HF RL, DeepSeek etc are central to this and extending all of this for the Neuron based system is key focussing on enabling large scale training. Experience is post-training strategies like DPO/PPO/HF torch-tune will additional strength and aligns with team success., Bachelor's degree in computer science or equivalent
-
5+ years of non-internship professional software development experience
-
5+ years of programming with at least one software programming language experience
-
5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
-
5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
-
Experience as a mentor, tech lead or leading an engineering team
-
Experience in machine learning, large scale training with LLMs and expertise in Pytorch.
Preferred Qualifications
-
Master's degree in computer science or equivalent
-
Experience in computer architecture
-
Previous software engineering expertise with Pytorch/Jax/Tensorflow, Distributed libraries and Frameworks, End-to-end Model Training.
Benefits & conditions
Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.
Inclusive Team Culture
Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.
Work/Life Balance
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.
Mentorship & Career Growth
We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.