> Markdown version of [/jobs/ext/2288806-cloud-customer-solutions-engineer-dc-gpu](https://www.wearedevelopers.com/jobs/ext/2288806-cloud-customer-solutions-engineer-dc-gpu). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud & Customer Solutions Engineer - DC GPU - **Company:** Advanced Micro Devices, Inc. - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Salary:** $204,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Computer Clusters, Nvidia CUDA, Databases, Computer Engineering, Software Debugging, Distributed Computing Environment, InfiniBand, Python (Programming Language), Open Source Technology, Prometheus, Datadog, Graphics Processing Unit (GPU), Google Cloud, Cloud Platform System, Performance Testing, Large Language Models, Grafana, Kubernetes, Information Technology, Production Code, Bare Metal, Slurm, Machine Learning Operations - **Published:** August 29, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88176267/1 ## About the Role You are a strong production engineer who is energized rather than drained by ambiguity, customer pressure, and environments you do not control. You can debug a distributed training hang at 2am, explain the root cause to a customer VP at 9am, and land the fix upstream by the end of the week. You measure success by customer production outcomes, not code merged or tickets closed. When something is broken on a cluster you touch, it is your problem until it is fixed or explicitly handed off., * 5+ years of production software or infrastructure engineering, including significant time operating or deploying systems in environments you did not build (level flexible for exceptional candidates) * Hands-on experience with GPU compute at scale: cluster deployment, distributed training or high-throughput inference, performance debugging, and workload optimization * Strong working knowledge of the modern AI infrastructure stack: Kubernetes and/or Slurm, containerized GPU workloads, collective communication libraries (RCCL/NCCL), high-performance networking (RoCE/InfiniBand), and observability tooling (Prometheus, Grafana) * Cloud platform depth (AWS, Azure, GCP, or NeoCloud environments), including hybrid and bare-metal deployment patterns * Proficiency in Python and at least one systems language; comfort navigating and modifying large codebases you did not write * Working familiarity with LLM application patterns - inference serving, RAG, and agentic workflows - sufficient to deploy and troubleshoot them in customer environments * Direct customer-facing experience: embedded deployments, technical escalations, on-site engagements, or equivalent * Experience with ROCm and AMD Instinct GPUs strongly preferred; deep CUDA-ecosystem experience with demonstrated ability to work cross-platform also valued * Open-source contribution history in AI/ML infrastructure projects is a plus PREFERRED ACADEMIC CREDENTIALS: * Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience ## Description As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers - frontier labs, NeoCloud providers, CSPs, and AI-native companies - to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets. To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production outcomes, and be the engineer in the room when things break at scale. What you learn in the field, you convert into durable improvements - to ROCm, to the open-source serving ecosystem, and to the reference architectures every subsequent deployment inherits., * Own customer deployments end-to-end: cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets * Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer-specific workloads and SLOs across cloud, NeoCloud, and bare-metal environments * Lead root-cause analysis and resolution of production incidents on customer clusters, including Sev-1 response, and drive fixes to permanent closure * Deploy agentic AI solutions into customer environments in partnership with Agentic Data Engineers, and own their production behavior within the engagement * Build the observability, benchmarking, and validation tooling needed to certify clusters as production-ready and keep them there * Transfer operational capability to customer teams - documentation, runbooks, and hands-on enablement - moving customers up the operator-autonomy ladder from assisted operation to independent production ownership * Contribute field learnings to the Applied AI team's skills library and engagement memory databases, so deployment knowledge compounds across the practice * Convert field findings into upstream contributions - ROCm issues and patches, serving-framework improvements, reference-architecture updates - and provide structured field signal to AMD product, software, and silicon teams, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)