Senior Cloud Engineer

Createfuturewe
Cardiff, UK
about 1 month ago
Apply on www.apply4u.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Cloud Engineering DevOps Github Knowledge-Based Systems Network Control Octopus Deploy Prometheus Datadog Large Language Models
+8 more
Grafana Multi-Agent Systems AI Platforms Gitlab-ci Kubernetes Machine Learning Operations Terraform Jenkins

Job description

production. You’ll help build and operate our Kubernetes-based platform for AI workloads, support CI/CD pipelines purpose-built for AI and agentic systems, and bring the SLOs, on-call, and incident response rigour that this space has historically lacked. Success looks like an AI platform that runs reliably and scales predictably.What you’ll be doing:Design, build, and maintain the Kubernetes infrastructure that supports AI workloads, including model serving, agent orchestration, and batch inference, with infrastructure as code in Terraform.Design and maintain CI/CD pipelines tailored to AI and agentic workflows, including model deployment, agent/tool updates, and prompt or configuration rollouts, with automated eval and regression checks before release.Define and track SLOs/SLAs for AI platform services, and bring SRE rigour to incident response, root cause analysis, and postmortems.Participate in on-call rotations and maintain clear, usable runbooks.Build observability for AI-specific

Requirements

concerns - latency, token usage, cost per request, model/agent error rates, and drift - with dashboards and alerting that surface issues before they reach users.Partner with the Inference Control Plane, Evals/Observability, and Context & Knowledge Platform teams so new agents, tools, and knowledge systems are built with operability in mind from day one.We’d love to talk to you if you:Have 4+ years’ experience in DevOps, SRE, or platform engineering, including production ownership of Kubernetes-based systems.Have hands-on experience operating AI or ML systems in production - model serving, LLM inference, or MLOps pipelines - a strong plus.Are strong in Python for automation and operational tooling, with production experience in Terraform and at least one major cloud provider (AWS or GCP).Have built and maintained CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Argo CD) and are comfortable with observability stacks (Datadog, Prometheus, Grafana).Have calm, rigorous incident management instincts, and ideally some familiarity with LLM/agent ecosystems (model APIs, vector databases, MCP, orchestration frameworks).What we’ll offer you:We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You’ll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed.We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.Our hiring processWe try to keep our hiring process clear, fair and respectful of your time. We aim to get back to everyone who applies and we will be upfront about where you are in the process.It usually looks like this:Call with our Talent Acquisition TeamRole specific capability interviewDepending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation. We will explain what is involved before anything happens.Inclusion at CreateFutureWe believe diverse teams build better workplaces and better products. We want CreateFuture to be a place where people feel able to be themselves and do their best work.If you need any adjustments or support during the application process, just. We will do what we can to help.We look forward to your application! #J-18808-Ljbffr

About the company

Working at CreateFutureCreateFuture is an AI-native consulting partner where people do work that matters and are supported to do it well. We work alongside organisations such as PayPal, adidas, NatWest, FanDuel and Money Saving Expert, building digital products and services that make a difference while always putting people first.We’re a team of creators. We write code, shape delivery, build go-to-market strategies, develop AI solutions and create the practices that support our people. We work side by side with our clients, challenging what’s not working and helping them to build the future. Our commitment to craft, quality, and culture has helped us scale to over 600 people in just a few years.35 days leave (including bank holidays).Private medical insurance.Enhanced parental and adoption leave.40 hours of paid learning and development.Join us on our journey. Let’s create tomorrow, together, today.About the role and team:You’ll be bringing SRE discipline to how AI platforms are run in

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all