> Markdown version of [/jobs/ext/2869069-principal-technical-account-manager-aws-enterprise-support-namer-sp](https://www.wearedevelopers.com/jobs/ext/2869069-principal-technical-account-manager-aws-enterprise-support-namer-sp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp - **Company:** Amazon.com, Inc. - **Location:** Austin, TX, United States - **Experience:** Experienced - **Salary:** $210,200.0 - $284,300.0 - **Contract:** Permanent contract - **Skills:** Mxnet, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Data Analysis, Artificial Neural Networks, Computer Vision, Cloud Computing, Computer Clusters, Databases, Distributed Computing Environment, Job Scheduling, Machine Learning, Performance Tuning, Tensorflow, SAS (Software), Pytorch, Large Language Models, Deep Learning, Parallel Computation, Information Technology, Performance Monitor, Slurm, Machine Learning Operations, GPT - **Published:** September 12, 2026 - **Apply:** https://dejobs.org/x/x/46A09F2C50144BD5A3184F110F336C50/job/ ## About the Role * 8+ years of working with Data & AI related technologies, including, but not limited to, AI/ML, GenAI, Analytics, Database, and/or Storage experience * 3+ years of hands-on experience designing, implementing, or consulting on large-scale ML training or inference architectures in a customer-facing role * 10+ years of IT development or implementation/consulting in the software, cloud computing, or AI/ML industries * Demonstrated ability to serve as a trusted technical advisor to enterprise customers, * Experience with deep learning libraries such as PyTorch, TensorFlow, MxNet Research publications in computer vision, deep learning or machine learning at peer-reviewed workshops, conferences or journals * Experience with training and deploying machine learning systems to solve large-scale optimizations * 5+ years of solving problems with technology in the Healthcare/Life Sciences Industry experience * Experience in the financial services industry ## Description Deliver Strategic Technical Engagements - Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 UltraServers), and multi-node NCCL communication tuning over EFA's SRD protocol. Architect and Validate Innovative Solutions - Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron-LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale.. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference. Enable Customer Success - Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads. Enable Business Critical Outcomes - Partner with with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node UltraClusters, and drive operational efficiency through proactive monitoring, automated failure recovery (HyperPod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share refrerence architecture, performance , and benchmarks with broader TAM and Technical communities Serve as Trusted Advisor and Advocate - Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC-to-AI convergence journey. A day in the life Your day will be dynamic and impactful, involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Trainium cluster deployments, and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders, architect innovative AI/ML implementations - from Slurm-managed PCS clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows - and provide expert guidance that bridges machine learning infrastructure with business objectives. You will partner with TAMs, SAs, and service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC-to-AI convergence patterns (simulation-surrogate loops, physics-informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) - [AI That Fits Your Business, Not the Other Way Around](https://www.wearedevelopers.com/videos/100148-ai-that-fits-your-business-not-the-other-way-around) - [Speak, Code, Deploy: Transforming Developer Experience with Voice Commands](https://www.wearedevelopers.com/videos/1159-speak-code-deploy-transforming-developer-experience-with-voice-commands) ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j)