> Markdown version of [/jobs/ext/2867530-principal-ai-ml-hpc-specialist-technical-account-manager-stam-aws-enterprise-support-namer-sp](https://www.wearedevelopers.com/jobs/ext/2867530-principal-ai-ml-hpc-specialist-technical-account-manager-stam-aws-enterprise-support-namer-sp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp - **Company:** Amazon.com, Inc. - **Location:** Austin, TX, United States - **Experience:** Experienced - **Salary:** $182,800.0 - $247,300.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Artificial Neural Networks, Cloud Computing, Computer Clusters, File Systems, Distributed Computing Environment, Distributed Systems, General Parallel File Systems, InfiniBand, Job Scheduling, Machine Learning, NetCDF, Performance Tuning, Tensorflow, Prometheus, SAS (Software), Graphics Processing Unit (GPU), Pytorch, Large Language Models, Grafana, Deep Learning, Parallel Computation, Information Technology, Performance Monitor, Slurm, Machine Learning Operations, Cloudwatch, GPT, Data Pipelines, Docker - **Published:** September 12, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28015340/Principal-Ai-Ml-Hpc-Specialist-Technical-Account-Manager-Stam-Aws-Enterprise-Support-Namer-Sp-Texas-Austin-7375 ## About the Role Bachelor's degree - 8+ years of experience in AI/ML, distributed computing, or GPU-accelerated infrastructure (e.g., model training, inference systems, HPC for ML) - 3+ years of hands-on experience designing, implementing, or consulting on large-scale ML training or inference architectures in a customer-facing role - 10+ years of IT development or implementation/consulting in the software, cloud computing, or AI/ML industries - Experience with at least one major deep learning framework (PyTorch, TensorFlow, JAX) in a production or research environment - Demonstrated ability to serve as a trusted technical advisor to enterprise customers, Deep experience with distributed training techniques including data parallelism, model parallelism, pipeline parallelism, and Fully Sharded Data Parallel (PyTorch FSDP) - Experience with distributed training frameworks such as PyTorch DDP, DeepSpeed, and Megatron-LM for multi-node model training - Hands-on experience with GPU/accelerator cluster infrastructure: NVIDIA Blackwell (GB200, B200, B300), H100/H200 GPUs, AWS Trainium (Trn3/Trn2), NVLink/NVSwitch, InfiniBand or Elastic Fabric Adapter (EFA), and NCCL collective communications tuning - Experience with AWS Neuron SDK (torch-neuronx, neuronx-nemo-megatron) for compiling and optimizing models on Trainium and Inferentia (Inf2) instances - Familiarity with SageMaker HyperPod for managed distributed training clusters including automated health checks, node replacement, and checkpoint-based recovery - Experience with HPC job schedulers (Slurm, PBS, LSF) for orchestrating multi-node ML training workloads - Experience with high-performance parallel file systems (Amazon FSx for Lustre, GPFS/Spectrum Scale) for ML data pipelines - Familiarity with AWS Parallel Computing Service (PCS), AWS ParallelCluster, AWS Batch, or equivalent managed HPC/ML cluster services - Experience training or fine-tuning large language models (LLMs) such as Llama, GPT, or similar transformer architectures at multi-billion parameter scale - Understanding of HPC-AI convergence patterns: simulation-surrogate loops, physics-informed neural networks (PINNs), graph neural networks for molecular property prediction, and data format interoperability (HDF5, VTK, NetCDF to ML-ready tensors) - Knowledge of ML Ops tooling, container orchestration for training (Docker, Enroot, Pyxis), and Deep Learning AMIs (DLAMIs) - Experience with cluster observability and monitoring for GPU/Trainium utilization, training throughput, and job performance (CloudWatch, Prometheus, Grafana) - Experience with EC2 Capacity Blocks for ML, Capacity Reservations, or similar GPU capacity planning strategies - Experience with pipeline orchestration using AWS Step Functions for simulation-ML workflows - Experience with containers, EKS and ECS - Track record of driving operational excellence and proactive risk mitigation for mission-critical AI/ML workloads - AWS certifications (Solutions Architect Professional, Machine Learning Specialty) preferred ## Description Deliver Strategic Technical Engagements - Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 UltraServers), and multi-node NCCL communication tuning over EFA's SRD protocol. Architect and Validate Innovative Solutions - Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron-LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale.. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference. Enable Customer Success - Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads. Enable Business Critical Outcomes - Partner with with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node UltraClusters, and drive operational efficiency through proactive monitoring, automated failure recovery (HyperPod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share refrerence architecture, performance , and benchmarks with broader TAM and Technical communities Serve as Trusted Advisor and Advocate - Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC-to-AI convergence journey. A day in the life Your day will be dynamic and impactful, involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Trainium cluster deployments, and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders, architect innovative AI/ML implementations - from Slurm-managed PCS clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows - and provide expert guidance that bridges machine learning infrastructure with business objectives. You will partner with TAMs, SAs, and service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC-to-AI convergence patterns (simulation-surrogate loops, physics-informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Optimizing your AI/ML workloads for sustainability](https://www.wearedevelopers.com/videos/570-optimizing-your-ai-ml-workloads-for-sustainability) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)