> Markdown version of [/jobs/ext/2118558-ai-systems-engineer-hpc](https://www.wearedevelopers.com/jobs/ext/2118558-ai-systems-engineer-hpc). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Systems Engineer - HPC - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Salary:** $173,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Ubuntu (Operating System), Computer Clusters, Computer Engineering, Data as a Services, Distributed Systems, Python (Programming Language), Kernel-Based Virtual Machine, Ansible, Prometheus, Azure Machine Learning, Web Services, AI Infrastructure, Graphics Processing Unit (GPU), High Performance Computing, Saltstack, Grafana, Backend, AI Platforms, Kubernetes, Information Technology, Slurm, Terraform - **Published:** August 19, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17984529?backUrl=%2Fcareer%2F17984529%2FAi-Systems-Engineer-Hpc-California-San-Jose ## About the Role You have a passion for learning. You are passionate about the field of large-scale distributed computing in AI and HPC workloads. You take responsibility for end-to-end outcomes of your efforts. You want to build scalable and highly performant HPC/AI/Data services with AMD hardware, software,peopleand processes. You have a curiosity to learn and improve scalable HPC systems. You havesignificant experiencein working across a globally distributed organization., * Experience in developingPython basedAI apps and UI * HPC infrastructure engineering for AI/HPC domain * SLURM and Kubernetes management * Managing GPU clustersoptimizingGPU-based services/tools/software * Experience in creating web services with HPC backend (like AI) * Proficiencyin RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and Cluster interconnect with 400G networking. * Demonstrated experience with AI workload schedulers and allocation optimization. * Automation/monitoring tool - Ansible / Saltstack, Terraform, Prometheus, Grafana * Strong organizational, problem-solving, and troubleshooting skills, with the ability to manage multiple projects simultaneously. * Excellent verbal and written communication skills, with the ability to collaborate effectively with team members and stakeholders at all levels of the organization. EDUCATION: Bachelor's or master's degree in computer science or computer engineering preferred. ## Description We are seeking an AI Systems Engineer to join our AMD ITcomputeplatforms engineering team. The AI Systems Engineeris responsible forthe design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers., * Develop, implement, andmaintainGPU-based clusters, ensuringoptimalperformance * Administer ML/AI platforms - Distributed ML services, LLMsandAIinferencing, by managing deployments, resource allocation, monitoring, and security. * Automate system provisioning and Cluster management end to end * Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technicalexpertise. * Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards. * Use AI/ML to continuously improve internal processes and tools that are used in end-to-end delivery of your services in this team, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [The Future of Computing: AI Technologies in the Exascale Era](https://www.wearedevelopers.com/videos/1106-the-future-of-computing-ai-technologies-in-the-exascale-era) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)