> Markdown version of [/jobs/ext/3030629-ai-infrastructure-operations-engineer](https://www.wearedevelopers.com/jobs/ext/3030629-ai-infrastructure-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Operations Engineer - **Company:** Accenture - **Location:** Denver, CO, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Computer Clusters, Configuration Management, Cyber Security, Data Security, Node.Js, Service Pack, Systems Architecture, Scripting, Graphics Processing Unit (GPU), Enterprise Software Applications, High Performance Computing, Containerization, Kubernetes, Information Technology, Bare Metal, Data Management - **Published:** September 22, 2026 - **Apply:** https://www.careerbuilder.com/job-details/ai-infrastructure-operations-engineer-denver-co--e900ee66-e7f2-417b-a4c7-83af9e2be3d2 ## About the Role Artificial Intelligence (AI), Automation, Benchmarking, Business Transformation, Capacity Management, Cloud Computing, Computer Systems, Configuration Management, Continuous Improvement, Cost Control, Ecosystems, Energy Efficiency, GPU (Graphics Processing Unit), Identify Issues, Incident Management, Incident Response, Information/Data Security (InfoSec), Operational Support, Operations Processes, Professional Services, Scripting (Scripting Languages), Service Delivery, Simulation, Software Patches, Support Documentation, System Architecture, Technical Leadership, Technical Operations, Validation Plan ## Description * Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements. * Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads. * Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls. * Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation. * Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management. * Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads. * Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation. * Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads, with emphasis on availability, resiliency, scalability, energy efficiency, and cost management. Travel may be required for this role. The amount of travel will vary from 25% to 60% depending on business need and client requirements. ## Related Videos - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [Navigating the AI Wave in DevOps](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)