> Markdown version of [/jobs/ext/82996-senior-sre-devops-engineer-ai-platform-engineering](https://www.wearedevelopers.com/jobs/ext/82996-senior-sre-devops-engineer-ai-platform-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior SRE / DevOps Engineer - AI Platform Engineering - **Company:** Create Digital Solutions - **Location:** London, UK - **Experience:** Expert - **Salary:** £60,000.0 - £80,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon Cloudfront, Amazon Elastic Compute Cloud, Cloud Engineering, Continuous Integration, DevOps, Distributed Systems, Identity and Access Management, Key Management, Routing, Open Source Technology, Software Architecture, Zero Trust Network Access, Systems Integration, Datadog, Data Logging, Delivery Pipeline, Large Language Models, Multi-Agent Systems, Amazon Virtual Private Cloud (VPC), Usage Tracking, Event Driven Architecture, Containerization, AI Platforms, Kubernetes, HuggingFace, Functional Programming, Api Gateway, Terraform - **Published:** May 16, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=305e133bb9d3c02a ## About the Role Do you have experience in Terraform?, Do you have a Master's degree?, This role requires commercial experience operating AI/LLM workloads in production. If your experience is primarily traditional DevOps/Cloud Engineering without hands-on exposure to: * LLMs / AI inference workloads * OpenAI, Bedrock, HuggingFace or hosted models * AI routing/orchestration * LLM observability and monitoring …then this role is unlikely to be suitable. We are specifically looking for engineers who have helped build or operate AI-enabled platforms at scale., AWS & Platform Engineering * Strong commercial AWS production experience * Hands-on experience with ECS, EC2, Lambda, CloudFront, IAM, VPC networking, and multi-account AWS environments * Strong troubleshooting and operational support experience * Experience supporting multiple teams/projects on shared platforms Terraform & DevOps * Strong Terraform experience including reusable module design and CI/CD integration * Container platform experience (ECS preferred) * Experience building shared/internal engineering platforms and reusable infrastructure patterns AI / LLM Experience (ESSENTIAL) You must have commercial experience managing, routing, hosting, or supporting AI/LLM workloads in production. Examples include: * OpenAI integrations * AWS Bedrock * HuggingFace * vLLM * Hosted inference services * AI orchestration/routing layers * Model-serving infrastructure You should also have familiarity with: * AI workload scaling and cost optimisation * GPU/compute-heavy workloads * LLM observability and monitoring * Prompt/API routing strategies * Agentic workflows and orchestration patterns Candidates without meaningful production AI/LLM experience will not be considered. Desirable Experience * GPU or specialised compute environments * Vector databases / embedding pipelines * Event-driven architectures * Platform engineering / Internal Developer Platforms * Security best practices (IAM, secrets management, Zero Trust) * Kubernetes/EKS exposure * AI gateways or AI proxy technologies Wider Skills * Comfortable making and defending architectural decisions * Strong communication and stakeholder management * Detail-oriented with a bias toward automation and standardisation * Comfortable operating across multiple concurrent workstreams ## Description We are looking for a Senior SRE / DevOps Engineer to help design and operate a shared AWS-based AI platform supporting multiple AI-driven products and services. A core objective is enabling teams to switch between AI providers and models (e.g. OpenAI, Anthropic, AWS Bedrock, hosted/open-source models) without impacting downstream applications or customer experience. You will help define reusable infrastructure patterns, deployment standards, observability tooling, and operational best practices that allow teams to rapidly deliver AI workloads while maintaining reliability, security, scalability, and cost control. This is a hands-on role suited to someone comfortable making architectural decisions, driving standardisation, and operating cloud-native AI systems in production., * Design and implement AI-agnostic platform patterns enabling interchangeable LLM/model providers * Build routing and integration patterns for hosted and third-party AI services * Support AI inference workloads and model-serving infrastructure in production * Implement observability, monitoring, usage tracking, latency analysis, and cost visibility for AI systems * Contribute to agentic workflow and orchestration capabilities * Define reusable AWS architecture patterns and infrastructure standards * Build and maintain Terraform modules and CI/CD deployment pipelines * Standardise deployment and operational practices across teams through automation and shared tooling * Embed SRE best practices including monitoring, logging, tracing, SLIs/SLOs, and alerting * Operate and troubleshoot distributed systems, API gateways, and container platforms ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)