> Markdown version of [/jobs/ext/1119248-ai-infrastructure-architecture-specialist](https://www.wearedevelopers.com/jobs/ext/1119248-ai-infrastructure-architecture-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Architecture Specialist - **Company:** Accenture - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, C++ (Programming Language), Computer Clusters, Code Review, Computer Engineering, System Configuration, Software Debugging, Interoperability, Python (Programming Language), Machine Learning, Workflow Management Systems, AI Infrastructure, Kubernetes, Information Technology, Deployment Automation, Machine Learning Operations, Data Pipelines, Docker, Programming Languages - **Published:** July 2, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=1b385a97948c319e ## About the Role * Bachelor's Degree in Computer Science, Computer Engineering, related Engineering field BASIC (REQUIRED) QUALIFICATION * Experienced with coding, building, monitoring, troubleshooting applications of AI/ML models; selecting, designing and infrastructure for deploying and running them on premise or on public cloud. * Strong understanding of AI and machine learning as a subject. * Strong understanding of computing infrastructure a subject, preferred knowledge of AI infrastructure. * Proficient in programming languages such as Python, Java, or C++. * Experience with data pipeline and workflow management tools (e.g., Apache Airflow, Kubeflow). * Strong problem-solving skills and ability to work in a fast-paced environment. * Excellent communication and collaboration skills. * Experience in AI/ML infrastructure engineering or related roles on a hyperscaler platform for deploying large scale solutions. ## Description As a hands-on Infrastructure Architect, you are an experienced engineer with several years in infrastructure engineering who now takes on more complex, higher-impact work designing and optimizing the AI and machine learning infrastructure that powers real-world applications. Working alongside senior architects and engineers - and increasingly leading your own workstreams - you apply proven skills in coding, testing, configuring, deploying, monitoring, and troubleshooting AI systems and the infrastructure they run on. Day to day, you architect and optimize infrastructure components, write and review code and deployment scripts, design and tune cloud and on-premises compute resources such as GPU clusters and distributed training environments, deploy AI systems and models into production, and build and optimize data pipelines that feed AI and ML workflows. You optimize the computational stack for performance, cost, power, and scalability, monitor AI systems and infrastructure health across both InfraOps and MLOps disciplines, perform AI monitoring to track model and system performance, and independently troubleshoot and resolve complex issues across the stack. You also mentor junior engineers, contribute to architectural decisions, and help establish best practices. This is a hands-on, ownership-driven where you apply and deepen your expertise across modern tools and platforms - including container orchestration, model serving, CI/CD pipelines, InfraOps, MLOps, and AI monitoring - while making meaningful contributions to infrastructure that enables AI-driven business outcomes. THE WORK * Write, review, and debug code, scripts, and infrastructure-as-code for AI infrastructure, automation, and tooling, setting standards for quality across the team. * Architect, configure, and provision compute resources across cloud and on-premises environments, including GPU clusters and distributed training setups, optimizing for performance and utilization. * Design and maintain deployment automation and CI/CD pipelines to support reliable, repeatable releases of AI systems, models, and applications. * Deploy AI systems, models, and data pipelines into production, defining and improving the processes and best practices others follow. * Lead container orchestration and model serving using tools such as Docker, Kubernetes, and model deployment frameworks. * Architect and optimize the computational stack for performance, power, cost, and scalability, balancing trade-offs against business goals. * Evaluate and select tools, frameworks, and platforms, making recommendations that shape the infrastructure roadmap. * Integrate AI models and systems into existing enterprise systems, ensuring interoperability, security, and regulatory compliance. * Own AI monitoring and infrastructure health across InfraOps and MLOps, tracking performance, reliability, and utilization, and driving remediation. * Independently troubleshoot and resolve complex issues across the computational stack - hardware, networking, software, and models - and lead root-cause analysis. * Mentor junior engineers and lead code reviews, providing technical direction and supporting their growth. * Define and document architecture standards, processes, and procedures, and apply security, cost-efficiency, and scalability best practices across the infrastructure. ## Related Videos - [Reference Architecture of AI in the Cloud](https://www.wearedevelopers.com/videos/1613-reference-architecture-of-ai-in-the-cloud) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [6 Reasons to Use Java For Your Next AI Project](https://www.wearedevelopers.com/magazine/111-6-reasons-to-use-java-for-your-next-ai-project) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)