> Markdown version of [/jobs/ext/2419517-ai-hpc-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2419517-ai-hpc-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI & HPC Infrastructure Engineer - **Company:** Accenture - **Location:** Beaverton, OR, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Cloud Computing, Cyber Security, Nvidia CUDA, Information Systems, Computer Networks, Data Security, Systems Architecture, AI Infrastructure, Graphics Processing Unit (GPU), System Availability, Large Language Models, Containerization, Kubernetes, Information Technology, Low Latency, Bare Metal, Data Management, Slurm, Machine Learning Operations, TensorRT, Virtual Agents, Data Pipelines - **Published:** August 21, 2026 - **Apply:** https://www.careerbuilder.com/job-details/ai-hpc-infrastructure-engineer-beaverton-or--7e23ea96-03e5-4681-bb3f-c9debddac597 ## About the Role Application Programming Interface (API), Artificial Intelligence (AI), Benchmarking, CPU (Central Processing Unit), CUDA (Compute Unified Device Architecture), Cloud Computing, Computer Networks, Cost Control, Data Management, Documentation, Ecosystems, Emerging Technology, Energy Efficiency, GPU (Graphics Processing Unit), High Availability, Identify Issues, Inference Engine, Information Technology & Information Systems, Information/Data Security (InfoSec), MCP - Microsoft Certified Professional, System Architecture, Technical Leadership, Use Cases ## Description * Design and implement AI infrastructure and accelerated computing solutions, aligning system architecture and deployment roadmaps to industry-specific performance, scalability, resiliency, and governance needs * Deploy, configure, and manage XPU-based clusters (GPU, DPU, LPU, CPU) across bare-metal and containerized environments using workload schedulers (Slurm, Run:ai), Kubernetes orchestration, and container platforms to deliver scalable AI infrastructure services including Bare-Metal-aaS, GPUaaS, AIaaS, Token-aaS, model serving, and agentic AI frameworks * Integrate AI infrastructure platforms with existing IT systems, data pipelines, security frameworks, model-serving endpoints, and enterprise governance controls * Design and implement agentic AI infrastructure by integrating platform services, model endpoints, tool and function calling, retrieval patterns, and workflow orchestration with observability, identity, and policy controls through secure, deterministic APIs to support governed enterprise use cases * Build and integrate MCP servers, tools, connectors, and adapters that allows agents to monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency * Architect and deploy with NVIDIA platform tools including Base Command Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang), inference orchestration (Triton Inference Server, NVIDIA Dynamo, llm-d), and GPU benchmarking and validation tools (MLPerf, NCCL tests, fio, iperf) to deploy, tune, profile, and validate AI cluster performance across compute and networking layers including multi-node training and inference workloads * Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks * Provide technical guidance, troubleshooting, and optimization across AI workloads including large-scale training, inference, multi-node simulations, and agentic pipelines while leveraging digital twins to validate infrastructure and drive performance, scalability, energy efficiency, and token cost optimization Travel may be required for this role. The amount of travel will vary from 25% to 100% depending on business need and client requirements. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)