> Markdown version of [/jobs/ext/1809750-ai-hpc-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1809750-ai-hpc-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI & HPC Infrastructure Engineer - **Company:** Accenture - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Salary:** $94,400.0 - $266,300.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Cloud Computing, Computer Clusters, Nvidia CUDA, Microprocessors, File Systems, Ethernet, InfiniBand, JSON, Python (Programming Language), Ansible, Tensorflow, Systems Architecture, Systems Integration, Weka, YAML, Openapi, AI Infrastructure, Jupyter Notebook, Graphics Processing Unit (GPU), Google Cloud, Pytorch, System Availability, Large Language Models, Generative AI, Data Center Networking, Build Management, Containerization, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Storage Technologies, Information Technology, Low Latency, Bare Metal, Slurm, Machine Learning Operations, TensorRT, Virtual Agents, Nutanix, Api Design, Restful APIs, Terraform, Webhooks, Data Pipelines, Automation Anywhere, Devsecops, Nvme, Vmware - **Published:** July 8, 2026 - **Apply:** https://dejobs.org/x/x/C80F0159F1644EDF8A44C9EFAE5C37F9/job/ ## About the Role * Minimum of 5+ years of experience designing, deploying, and managing AI infrastructure and accelerated computing environments across on-premises, cloud, and hybrid environments for hyperscaler, neocloud, large enterprise, Telco/Mobile, Financial Services, Life Sciences, Manufacturing, and/or Retail clients. * Minimum of 5+ years of hands-on experience with accelerated computing platforms, including GPUs, DPUs, LPUs, CPUs, high-speed interconnects such as InfiniBand or Ethernet, data center networking such as SONiC, and AI storage architectures including NVMe, NVMe-oF, parallel file systems, VAST, Weka, or DDN. * Minimum of 5+ years of experience with cluster management, workload scheduling, orchestration, observability, and infrastructure automation using platforms and tools such as Kubernetes, Slurm, Run:ai, AWS, Azure, GCP, VMware, Nutanix, Python, Terraform, and Ansible. * Bachelor's degree or equivalent (minimum 12 years) work experience. If Associate's Degree, must have minimum 6 years work experience. Preferred Skills and Qualifications: * 2+ years of experience implementing MLOps, LLMOps, agentic AI, and DevSecOps frameworks to enable secure, automated, governed, and reproducible AI workflows. * 2+ years of experience developing APIs, integration services, automation workflows, or platform services using Python and modern API patterns such as REST, OpenAPI, JSON/YAML schemas, webhooks, and event-driven integrations. * Experience designing and implementing agentic AI infrastructure, including LLM inference, tool/function calling, retrieval-augmented generation (RAG), agent orchestration, secure API integration, policy-based governance, and deterministic platform APIs. * Experience building and integrating MCP servers, tools, connectors, and adapters that allow agents to monitor, troubleshoot, and tune infrastructure for high availability, low-latency networking, workload resiliency, and intelligent observability. * Experience using NVIDIA platform tools including Base Command Manager (BCM), NGC, NCCL, NVLink, CUDA, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, llm-d, vLLM, SGLang, MLPerf, NCCL tests, fio, and iperf to deploy, tune, profile, and validate AI cluster performance. * Experience managing the deployment of 1,000+ GPU clusters for AI, HPC, and agentic AI workloads with infrastructure services enabled. * Design and build experience in AI Cloud platforms from CoreWeave, Nebius, and other specialty cloud providers. * Knowledge of machine learning and AI frameworks such as TensorFlow, PyTorch, JAX, Jupyter notebooks, and Google Colab environments. * Industry certifications in NVIDIA infrastructure, public cloud providers, data science, infrastructure automation, networking, or security are a plus. ## Description The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered by AI, accelerated computing, and high-performance workloads. We bring together deep technical expertise across cloud, on-premises, and hybrid environments to design, build, and operate advanced infrastructure that powers AI platforms, GPU-accelerated workloads, large-scale models, simulations, and emerging agentic AI solutions at scale. Our solutions enable some of our most strategic and mission-critical clients to unlock new levels of performance, efficiency, governance, and innovation. Our remit spans the full lifecycle-from strategy and architecture through implementation and operations-driving modernization across the entire infrastructure stack. We collaborate across the ecosystem to harness emerging technologies, fuel growth, and transform industries. In this rapidly growing market, our team is leading the way in shaping how enterprises leverage AI infrastructure to drive breakthrough innovation and reimagine what is possible., * Design and implement AI infrastructure and accelerated computing solutions, aligning system architecture and deployment roadmaps to industry-specific performance, scalability, resiliency, and governance needs * Deploy, configure, and manage XPU-based clusters (GPU, DPU, LPU, CPU) across bare-metal and containerized environments using workload schedulers (Slurm, Run:ai), Kubernetes orchestration, and container platforms to deliver scalable AI infrastructure services including Bare-Metal-aaS, GPUaaS, AIaaS, Token-aaS, model serving, and agentic AI frameworks * Integrate AI infrastructure platforms with existing IT systems, data pipelines, security frameworks, model-serving endpoints, and enterprise governance controls * Design and implement agentic AI infrastructure by integrating platform services, model endpoints, tool and function calling, retrieval patterns, and workflow orchestration with observability, identity, and policy controls through secure, deterministic APIs to support governed enterprise use cases * Build and integrate MCP servers, tools, connectors, and adapters that allows agents to monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency * Architect and deploy with NVIDIA platform tools including Base Command Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang), inference orchestration (Triton Inference Server, NVIDIA Dynamo, llm-d), and GPU benchmarking and validation tools (MLPerf, NCCL tests, fio, iperf) to deploy, tune, profile, and validate AI cluster performance across compute and networking layers including multi-node training and inference workloads * Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks * Provide technical guidance, troubleshooting, and optimization across AI workloads including large-scale training, inference, multi-node simulations, and agentic pipelines while leveraging digital twins to validate infrastructure and drive performance, scalability, energy efficiency, and token cost optimization Travel may be required for this role. The amount of travel will vary from 25% to 100% depending on business need and client requirements. ## Related Videos - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)