AI Infrastructure Lead Architect

Accenture
London, UK
14 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£50,000.0 - £75,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Microsoft Azure C++ (Programming Language) Cloud Computing Computer Clusters Computer Engineering Continuous Integration Python (Programming Language) Machine Learning
+8 more
Azure Machine Learning Software Technical Review Workflow Management Systems AI Infrastructure Kubernetes Information Technology Machine Learning Operations Data Pipelines

Job description

  • Own the end-to-end architecture and design of optimized compute infrastructure for large-scale AI/ML systems, including large-scale distributed training environments, from concept through delivery
  • Develop and evaluate architecture alternatives, weighing trade-offs across compute, networking, storage, orchestration, and model serving to make rational, well-justified decisions tailored to each clients situation and standards
  • Lead architecture assessments and reviews of existing and proposed environments, identifying gaps, risks, bottlenecks, and optimization opportunities, and recommending remediation
  • Drive architectural decision-making, documenting rationale, trade-offs, and assumptions so decisions are transparent, defensible, and aligned with business SLAs and standards
  • Define and maintain the AI infrastructure roadmap, planning capacity, scaling, and technology evolution in step with business and product goals
  • Architect and optimize the full computational stack for performance, power, cost, and scalability, ensuring infrastructure meets business SLAs while being deliberately engineered for cost-efficiency
  • Design and tune large-scale GPU clusters and distributed training systems, including accelerator selection, interconnect/networking, and storage for high-throughput training workloads
  • Serve as the authoritative AI infrastructure expert in at least one hyperscaler cloud (AWS, Azure, or GCP), applying deep knowledge of its AI/ML services, accelerators, networking, and cost levers
  • Design deployment, automation, and CI/CD strategies for reliable, repeatable, and scalable releases of AI systems, models, and data pipelines into production
  • Establish AI monitoring and observability strategy across InfraOps and MLOps, defining SLAs, SLOs, alerting, and performance/cost tracking, and driving continuous optimization
  • Integrate AI/ML systems into enterprise environments, ensuring interoperability, security, compliance, and adherence to regulatory and client standards
  • Lead capacity planning and cost modeling, forecasting compute needs and engineering cost-efficiency into the architecture without compromising performance
  • Collaborate with clients, stakeholders, and engineering teams to align infrastructure decisions with business outcomes, translating requirements into actionable architecture and standards
  • Set technical direction, standards, and best practices, mentoring engineers and architects and leading design and code reviews across the team

Technologies:

  • AI
  • Airflow
  • AWS
  • Architect
  • Azure
  • CI/CD
  • Cloud
  • GCP
  • Java
  • Kubeflow
  • Machine Learning
  • MLOps
  • Model Serving
  • Python
  • Security

Requirements

  • Bachelors degree in Computer Science, Computer Engineering, or a related Engineering field
  • Solid background in coding, building, monitoring, and troubleshooting AI/ML model applications, including selecting, designing, and implementing infrastructure for deployment on premises or in public cloud
  • Strong understanding of AI and machine learning
  • Strong understanding of computing infrastructure, with preferred knowledge of AI infrastructure
  • Good proficiency in programming languages such as Python, Java, or C++
  • Experience with data pipeline and workflow management tools such as Apache Airflow or Kubeflow
  • Strong problem-solving skills and ability to work in a fast-paced environment
  • Excellent communication and collaboration skills
  • Significant experience in AI/ML infrastructure engineering or related roles on a hyperscaler platform for deploying large-scale solutions
  • Proven experience leading and managing AI projects and teams
  • Strong project management skills, including the ability to manage multiple projects simultaneously
  • Demonstrated experience evaluating and selecting AI technologies and frameworks
  • Ability to work with cross-functional teams and drive project alignment

About the company

We are Accenture, a leading global professional services company helping the worlds leading businesses, governments, and organizations build their digital core, optimize operations, accelerate revenue growth, and enhance citizen services. We are a talent- and innovation-led company with approximately 791,000 people serving clients in more than 120 countries. Technology is at the core of our work, and we are a leader in cloud, data, and AI with strong ecosystem relationships and global delivery capability. This is a full-time role based in London, Paris, Berlin, or Madrid (Castellana 85). We are committed to shared success, creating 360 value for our clients, our people, our shareholders, partners, and communities, and we value diversity and equal opportunity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell · Coffee With Developers

Videos

See all

Related articles

See all