Senior DevOps Engineer (Infrastructure & MLOps)

Talenzon group
London, UK
10 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Databases Continuous Integration DevOps Github Python (Programming Language) NoSQL Ansible
+21 more
Prometheus Azure Machine Learning SQL Databases Datadog Data Logging Scripting Cloud Platform System Delivery Pipeline Grafana Gitlab Cloudformation Containerization Kubernetes Infrastructure Automation Frameworks Google Bigquery AWS Fargate Machine Learning Operations Terraform Software Version Control Data Pipelines Docker

Job description

Senior DevOps Engineer (Infrastructure & MLOps),February 15, 2026### Job DescriptionLocation: London, UK Work Model: On-site Role Type: Full-TimeWe are looking for a Senior DevOps Engineer (Infrastructure & MLOps) with strong experience in cloud infrastructure and CI/CD automation to join our client’s on-site team in London.This role focuses on designing, implementing, and managing highly available cloud infrastructure while supporting AI/ML platforms and production-grade deployment pipelines. You will work closely with engineering and data teams to deliver scalable, secure, and reliable systems.—### What You’ll Do* Design, implement, and manage highly available infrastructure for cloud-based platforms using Amazon Web Services* Architect and support AI/ML infrastructure including Lambda-based workloads and managed ML environments for training, hosting, and inference* Create and automate robust deployment pipelines using CI/CD tools such as GitHub Actions

Requirements

and GitLab* Build, maintain, and scale containerised applications using Docker and Kubernetes / ECS / Fargate* Implement MLOps best practices to streamline model transition from development to production* Ensure system scalability and reliability through monitoring, logging, and automated alerting using tools such as Datadog, Prometheus, Grafana, and MLflow* Collaborate with Product Engineers and Data Scientists to optimise performance, security, and infrastructure costs* Manage and evolve Infrastructure-as-Code environments—### What We’re Looking For#### Required Skills & Experience* 5+ years of experience in a DevOps, SRE, or infrastructure engineering role* Expert knowledge of cloud platforms (AWS preferred; GCP or Azure also valuable)* Strong experience with containerisation technologies (Docker, ECS, Kubernetes)* Proven experience designing and managing complex CI/CD pipelines* Experience with MLOps workflows (model versioning, retraining pipelines, feature stores)* Hands-on experience with monitoring and logging platforms* Strong scripting skills (Python required; Bash, Go, or similar beneficial)* Experience with Infrastructure-as-Code tools (Terraform, Ansible, or CloudFormation)* Excellent communication skills and ability to collaborate with both DevOps and Data Science teams on-site—#### Nice to Have* Direct experience with AWS Lambda and managed ML services* Experience with database systems (SQL/NoSQL) and orchestration of data pipelines including Google BigQuery* Knowledge of security and data privacy best practices in regulated environments (e.g. healthcare / sensitive data)* Previous experience working in fast-paced startup or scale-up environments—Location: London, UK Work Model: On-site Role Type: Full-TimeLocation,Experience levelMid-Senior level## Work Location #J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:44 min

Defining core roles and responsibilities in MLOps teams

Bas Geerdink · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · World Congress 2023

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all