> Markdown version of [/jobs/ext/2696696-ai-infrastructure-operations-engineer](https://www.wearedevelopers.com/jobs/ext/2696696-ai-infrastructure-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Operations Engineer - **Company:** Accenture - **Location:** Arlington, VA, United States - **Salary:** $94,400.0 - $266,300.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Bash Shell, Cloud Computing, Computer Clusters, Configuration Management, Nvidia CUDA, Microprocessors, File Systems, JSON, Python (Programming Language), Node.Js, Performance Tuning, Reliability Engineering, Ansible, Systems Architecture, YAML, Openapi, AI Infrastructure, Graphics Processing Unit (GPU), Enterprise Software Applications, High Performance Computing, Performance Testing, Large Language Models, Build Management, Containerization, Kubernetes, Infrastructure Automation Frameworks, Storage Technologies, Bare Metal, Data Management, Slurm, Restful APIs, Terraform, Webhooks, Nvme - **Published:** September 3, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3375502868&tx=VT7368THD&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Minimum of 5+ years of experience designing, deploying, and managing accelerated-computing infrastructure across on-premises, cloud, and hybrid environments for hyperscaler, neocloud, large enterprise, telecommunications, financial services, manufacturing, and/or retail clients. * Minimum of 5+ years of hands-on experience with accelerated-computing platforms, including GPUs, DPUs, and CPUs, high-bandwidth network fabrics, and AI based storage architectures such as parallel file systems, NVMe-oF, etc. * Minimum of 5+ years of experience with cluster management, workload scheduling, orchestration, observability, and infrastructure automation, including building operational tools and automation workflows with platforms such as Kubernetes, Slurm, Run:ai * Minimum 6 months hands-on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting. * Bachelor's degree or equivalent (minimum 12 years) work experience. (If Associate's Degree, must have minimum 6 years work experience), * Experience building AI infrastructure automation and operations tools, AgenticOps practices that enable secure, automated, governed, and reproducible platform operations. * Experience developing reusable infrastructure code leveraging Python, platform services, and automation workflows using REST APIs, OpenAPI, JSON/YAML schemas, webhooks, and event-driven integrations. * Experience operating large-scale GPU clusters, including capacity management, reliability engineering, change management, and performance validation for AI training, inference, HPC, and enterprise compute workloads. * Experience using NVIDIA platform tools and libraries including Base Command Manager (BCM), NGC, NCCL, CUDA-X, NVAIE, Dynamo, benchmarking tools to deploy, tune, profile, and validate cluster performance for training and inference workloads. * Experience managing deployments of 1,000+ GPU clusters with infrastructure services enabled for AI training, inference, high-performance, and enterprise compute environments. * Design and build experience in running LLMs across AI Cloud platforms from CoreWeave, Nebius, and other specialty providers * Knowledge of model deployments, tuning, and troubleshooting performance on AI Infrastructure * Industry certifications in accelerated-computing infrastructure, public cloud providers, infrastructure automation, networking, or security are a plus. ## Description * Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements. * Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads. * Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls. * Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation. * Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management. * Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads. * Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation. * Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads, with emphasis on availability, resiliency, scalability, energy efficiency, and cost management. Travel may be required for this role. The amount of travel will vary from 25% to 60% depending on business need and client requirements. ## Related Videos - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Reference Architecture of AI in the Cloud](https://www.wearedevelopers.com/videos/1613-reference-architecture-of-ai-in-the-cloud) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)