Senior SRE / DevOps Engineer - AI Platform Engineering

Create Digital Solutions
London, UK
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£60,000.0 - £80,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Cloudfront Amazon Elastic Compute Cloud Cloud Engineering Continuous Integration DevOps Distributed Systems Identity and Access Management Key Management Routing
+19 more
Open Source Technology Software Architecture Zero Trust Network Access Systems Integration Datadog Data Logging Delivery Pipeline Large Language Models Multi-Agent Systems Amazon Virtual Private Cloud (VPC) Usage Tracking Event Driven Architecture Containerization AI Platforms Kubernetes HuggingFace Functional Programming Api Gateway Terraform

Job description

We are looking for a Senior SRE / DevOps Engineer to help design and operate a shared AWS-based AI platform supporting multiple AI-driven products and services. A core objective is enabling teams to switch between AI providers and models (e.g. OpenAI, Anthropic, AWS Bedrock, hosted/open-source models) without impacting downstream applications or customer experience. You will help define reusable infrastructure patterns, deployment standards, observability tooling, and operational best practices that allow teams to rapidly deliver AI workloads while maintaining reliability, security, scalability, and cost control. This is a hands-on role suited to someone comfortable making architectural decisions, driving standardisation, and operating cloud-native AI systems in production., * Design and implement AI-agnostic platform patterns enabling interchangeable LLM/model providers

  • Build routing and integration patterns for hosted and third-party AI services
  • Support AI inference workloads and model-serving infrastructure in production
  • Implement observability, monitoring, usage tracking, latency analysis, and cost visibility for AI systems
  • Contribute to agentic workflow and orchestration capabilities
  • Define reusable AWS architecture patterns and infrastructure standards
  • Build and maintain Terraform modules and CI/CD deployment pipelines
  • Standardise deployment and operational practices across teams through automation and shared tooling
  • Embed SRE best practices including monitoring, logging, tracing, SLIs/SLOs, and alerting
  • Operate and troubleshoot distributed systems, API gateways, and container platforms

Requirements

Do you have experience in Terraform?, Do you have a Master’s degree?, This role requires commercial experience operating AI/LLM workloads in production. If your experience is primarily traditional DevOps/Cloud Engineering without hands-on exposure to:

  • LLMs / AI inference workloads
  • OpenAI, Bedrock, HuggingFace or hosted models
  • AI routing/orchestration
  • LLM observability and monitoring

…then this role is unlikely to be suitable.

We are specifically looking for engineers who have helped build or operate AI-enabled platforms at scale., AWS & Platform Engineering

  • Strong commercial AWS production experience
  • Hands-on experience with ECS, EC2, Lambda, CloudFront, IAM, VPC networking, and multi-account AWS environments
  • Strong troubleshooting and operational support experience
  • Experience supporting multiple teams/projects on shared platforms

Terraform & DevOps

  • Strong Terraform experience including reusable module design and CI/CD integration
  • Container platform experience (ECS preferred)
  • Experience building shared/internal engineering platforms and reusable infrastructure patterns

AI / LLM Experience (ESSENTIAL) You must have commercial experience managing, routing, hosting, or supporting AI/LLM workloads in production. Examples include:

  • OpenAI integrations
  • AWS Bedrock
  • HuggingFace
  • vLLM
  • Hosted inference services
  • AI orchestration/routing layers
  • Model-serving infrastructure

You should also have familiarity with:

  • AI workload scaling and cost optimisation
  • GPU/compute-heavy workloads
  • LLM observability and monitoring
  • Prompt/API routing strategies
  • Agentic workflows and orchestration patterns

Candidates without meaningful production AI/LLM experience will not be considered.

Desirable Experience

  • GPU or specialised compute environments
  • Vector databases / embedding pipelines
  • Event-driven architectures
  • Platform engineering / Internal Developer Platforms
  • Security best practices (IAM, secrets management, Zero Trust)
  • Kubernetes/EKS exposure
  • AI gateways or AI proxy technologies

Wider Skills

  • Comfortable making and defending architectural decisions
  • Strong communication and stakeholder management
  • Detail-oriented with a bias toward automation and standardisation
  • Comfortable operating across multiple concurrent workstreams

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

Videos

See all

Related articles

See all