> Markdown version of [/jobs/ext/2464360-senior-cloud-engineer](https://www.wearedevelopers.com/jobs/ext/2464360-senior-cloud-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Cloud Engineer - **Company:** CreateFuture - **Location:** Edinburgh, UK (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Cloud Engineering, DevOps, Github, Knowledge-Based Systems, Network Control, Octopus Deploy, Prometheus, Datadog, Large Language Models, Grafana, Multi-Agent Systems, AI Platforms, Gitlab-ci, Kubernetes, Low Latency, Machine Learning Operations, Terraform, Jenkins - **Published:** August 4, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=5d5f2fda88c105b4 ## About the Role * Have 4+ years' experience in DevOps, SRE, or platform engineering, including production ownership of Kubernetes-based systems. * Have hands-on experience operating AI or ML systems in production - model serving, LLM inference, or MLOps pipelines - a strong plus. * Are strong in Python for automation and operational tooling, with production experience in Terraform and at least one major cloud provider (AWS or GCP). * Have built and maintained CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Argo CD) and are comfortable with observability stacks (Datadog, Prometheus, Grafana). * Have calm, rigorous incident management instincts, and ideally some familiarity with LLM/agent ecosystems (model APIs, vector databases, MCP, orchestration frameworks)., Depending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation. We will explain what is involved before anything happens. ## Description You'll be bringing SRE discipline to how AI platforms are run in production. You'll help build and operate our Kubernetes-based platform for AI workloads, support CI/CD pipelines purpose-built for AI and agentic systems, and bring the SLOs, on-call, and incident response rigour that this space has historically lacked. Success looks like an AI platform that runs reliably and scales predictably., * Design, build, and maintain the Kubernetes infrastructure that supports AI workloads, including model serving, agent orchestration, and batch inference, with infrastructure as code in Terraform. * Design and maintain CI/CD pipelines tailored to AI and agentic workflows, including model deployment, agent/tool updates, and prompt or configuration rollouts, with automated eval and regression checks before release. * Define and track SLOs/SLAs for AI platform services, and bring SRE rigour to incident response, root cause analysis, and postmortems. * Participate in on-call rotations and maintain clear, usable runbooks. * Build observability for AI-specific concerns - latency, token usage, cost per request, model/agent error rates, and drift - with dashboards and alerting that surface issues before they reach users. * Partner with the Inference Control Plane, Evals/Observability, and Context & Knowledge Platform teams so new agents, tools, and knowledge systems are built with operability in mind from day one., We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe)