AI Infrastructure Operations Engineer

Accenture
Denver, CO, United States
3 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Cloud Computing Computer Clusters Configuration Management Cyber Security Data Security Node.Js Service Pack Systems Architecture Scripting Graphics Processing Unit (GPU) Enterprise Software Applications
+6 more
High Performance Computing Containerization Kubernetes Information Technology Bare Metal Data Management

Job description

  • Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements.
  • Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads.
  • Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls.
  • Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation.
  • Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management.
  • Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads.
  • Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation.
  • Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads, with emphasis on availability, resiliency, scalability, energy efficiency, and cost management.

Travel may be required for this role. The amount of travel will vary from 25% to 60% depending on business need and client requirements.

Requirements

Artificial Intelligence (AI), Automation, Benchmarking, Business Transformation, Capacity Management, Cloud Computing, Computer Systems, Configuration Management, Continuous Improvement, Cost Control, Ecosystems, Energy Efficiency, GPU (Graphics Processing Unit), Identify Issues, Incident Management, Incident Response, Information/Data Security (InfoSec), Operational Support, Operations Processes, Professional Services, Scripting (Scripting Languages), Service Delivery, Simulation, Software Patches, Support Documentation, System Architecture, Technical Leadership, Technical Operations, Validation Plan

About the company

Accenture is a global professional services company with leading capabilities in digital, cloud and security. Combining unmatched experience and specialized skills across more than 40 industries, we offer Strategy and, Interactive, Technology, and Operations services, all powered by the world’s largest network of Advanced Technology and Intelligent Operations centers. Our 738,000 people deliver on the promise of technology and human ingenuity every day, serving clients in more than 120 countries. We embrace the power of change to create value and shared success for our clients, people, shareholders, partners, and communities. Visit us at www.accenture.com.

The Global AI Infrastructure team enables resilient, high-performance compute environments for strategic clients across cloud, on-premises, and hybrid deployments. We design, build, and operate large-scale GPU and accelerated-computing infrastructure that supports demanding AI training and inference, simulation, and high-performance compute workloads. Our work spans strategy, architecture, modernization, operations, governance, and continuous improvement across the infrastructure stack. We build reusable operational tools, automation workflows, and platform capabilities that make repeatable infrastructure tasks safer, faster, and more scalable. We collaborate across the technology ecosystem to harness new capabilities, drive business transformation, and deliver dependable services at scale.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:52 min

Delivering diverse consulting services from agile to artificial intelligence

Ranjit Lopez Ranjit Lopez · Europe 2026 Virtual

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

1:11 min

Running high-performance edge computing on bare metal servers

Josip Stuhli Josip Stuhli · Coffee With Developers

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell · Coffee With Developers

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · World Congress 2023

Videos

See all

Related articles

See all