Senior DevOps & Systems Engineer

Centric Talent
Blackburn, UK
5 days ago
Apply on jobs4a.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Compensation
£60,000.0 - £65,000.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Apache HTTP Server Application Layers Systems Engineering Cloud Computing Configuration Management Databases System Configuration Continuous Integration Linux DevOps Disaster Recovery
+47 more
File Systems Domain Name System (DNS) Monitoring of Systems Web Servers Information Technology Operations Virtual Private Networks (VPN) Linux System Administration Linux Servers Microsoft Operating Systems Microsoft Servers Microsoft SQL Server Windows Servers MySQL Network Configuration and Change Management Nginx Package Management Systems Redis Ansible SAP (Applications) Server Administration Software Deployment Software Engineering TCP/IP Virtual Machines Software Vulnerability Management Data Logging Network Storage Transport Layer Security Enterprise Software Applications Load Balancing System Availability Delivery Pipeline Caching Database Performance Git SAPBasis Containerization Kubernetes Infrastructure Automation Frameworks Deployment Automation Patch Management Database Backup Terraform Network Server Software Version Control Docker Server Operating Systems & Platforms

Job description

The Senior DevOps & Systems Engineer will be responsible for the administration, reliability, security and continuous improvement of the TW Group infrastructure estate across both Linux and Microsoft environments.

This is a hands-on senior role for an experienced infrastructure engineer who is comfortable taking ownership of technical problems, investigating unfamiliar systems and delivering practical solutions with minimal supervision.

The successful candidate will work closely with the Software Development Team, IT Operations and selected third-party technology partners to support business-critical infrastructure, development environments, databases, container platforms and deployment pipelines.

The role will be expected to take ownership of day-to-day infrastructure operations while also improving automation, monitoring, resilience, documentation and disaster recovery capability across the wider IT estate., Administration, maintenance and improvement of Linux and Microsoft server environments.

  • Provisioning and configuration of new virtual machines, servers, services and supporting infrastructure.

  • Deployment and configuration of applications and infrastructure components.

  • Monitoring the health, performance, availability and security of production and internal systems.

  • Investigating and resolving infrastructure incidents, performance issues and system failures.

  • Designing and implementing appropriate monitoring, alerting and logging for new and existing systems.

  • Managing and improving configuration automation using Ansible.

  • Supporting and administering containerised environments using Docker, CRI-O and Kubernetes.

  • Managing Kubernetes clusters, workloads, services and associated infrastructure.

  • Administration and operational support of MySQL, Galera Cluster and Microsoft SQL Server environments.

  • 1 Building, maintaining and improving CI/CD pipelines.

  • Supporting software deployment and release processes in collaboration with the Development Team.

  • Maintaining system patching, upgrades, security configuration and vulnerability remediation.

  • Managing backup, recovery and disaster recovery processes.

  • Improving infrastructure resilience and reducing key-person dependencies through documentation and cross-training.

  • Reviewing existing infrastructure and recommending improvements where systems can be made more reliable, secure, maintainable or efficient.

  • Supporting infrastructure-related security incidents and remediation activities.

  • Maintaining clear technical documentation covering systems, infrastructure, recovery procedures and operational processes.

  • Working with third-party infrastructure and technology partners where specialist support is required.

  • Supporting the broader IT function where infrastructure, systems or operational responsibilities overlap.

Requirements

5+ years of professional experience in DevOps, Systems Administration, Infrastructure Engineering, Platform Engineering or a similar role.

  • Significant hands-on experience operating production-grade infrastructure.

  • Strong Linux administration experience.

  • Practical experience across both Linux and Microsoft server estates.

  • Proven experience managing infrastructure supporting business-critical applications.

  • Experience taking ownership of complex technical problems with limited supervision.

  • Proven ability to investigate unfamiliar systems and determine appropriate solutions.

  • Experience improving infrastructure reliability, automation, security or operational efficiency.

  • Previous experience operating in a senior or highly autonomous engineering role

Skills and Experience - Core Linux & Systems Administration

  • Strong professional experience administering Linux server environments in production.

  • Strong understanding of Linux system administration, including: users and permissions system services package management networking storage and filesystems scheduled tasks logging performance management troubleshooting.

  • Experience provisioning and configuring servers and virtual machines from initial deployment through to production readiness.

  • Strong understanding of production infrastructure reliability, availability and operational support.

  • Experience with system patching, upgrades and lifecycle management.

  • Experience diagnosing complex infrastructure issues independently.

Monitoring, Reliability & Security

  • Strong experience with system, application and security monitoring.

  • Experience designing and implementing monitoring and alerting for new infrastructure and applications.

  • Ability to identify meaningful service health indicators and create appropriate alerts.

  • Strong understanding of: infrastructure security access control patch management vulnerability remediation system hardening logging auditing certificate management
  • Experience responding to production incidents and security-related infrastructure issues.

  • Strong understanding of backup, recovery and disaster recovery principles.

  • Experience documenting and testing disaster recovery procedures.

Automation & Configuration Management

  • Strong working knowledge of Ansible.

Experience using automation to manage: server configuration 3.

  • Application deployment package installation security configuration environment setup recurring operational tasks
  • Ability to design reusable, maintainable and well-documented automation.

  • Strong understanding of infrastructure automation and configuration management principles.

  • Experience with Infrastructure as Code concepts.

Containers & Kubernetes

  • Strong experience with Docker and containerised workloads.

  • Experience with CRI-O, containerd or comparable container runtimes.

  • Practical experience administering Kubernetes clusters in production or business-critical environments.

  • Strong understanding of: deployments pods services ingress persistent storage secrets configuration resource allocation cluster networking availability monitoring troubleshooting
  • Ability to independently deploy, troubleshoot and maintain applications within Kubernetes.

  • Experience supporting highly available containerised environments

Database Administration MySQL / Galera

  • Strong experience administering MySQL environments.

  • Experience managing Galera Cluster or comparable highly available MySQL architectures.

  • Good understanding of: replication backup and restore performance
  • monitoring database troubleshooting permissions and access upgrades clustering high availability failover disaster recovery
  • Ability to investigate and resolve database performance and availability issues

Microsoft SQL Server

  • Experience administering Microsoft SQL Server / MSSQL environments.

  • Good understanding of: database backup and recovery
  • permissions maintenance
  • monitoring performance
  • troubleshooting availability resilience

Microsoft Systems Administration

  • Strong experience administering Microsoft Windows Server environments.

  • Experience installing, configuring and maintaining Windows-based servers and applications.

  • Experience with system monitoring, patching, security and operational support.

  • Strong understanding of Windows Server permissions, services and network configuration.
  • Ability to troubleshoot Windows infrastructure and application issues independently.

  • Experience supporting business-critical Microsoft workloads.

CI/CD & Development Operations

  • Strong experience with CI/CD pipelines.

  • Experience designing, maintaining and troubleshooting automated deployment pipelines.

  • Understanding of build, test and production deployment processes.

  • Strong working knowledge of Git and modern source control workflows.

  • Experience supporting software developers with infrastructure and deployment requirements.

  • Ability to improve deployment reliability through automation and standardisation.

  • Experience supporting multiple development, staging and production environments

Infrastructure & Networking

  • Strong understanding of: TCP/IP DNS HTTP / HTTPS firewalls VPNs load balancing reverse proxies TLS / SSL certificates
  • Experience administering web servers such as Nginx and Apache.

  • Understanding of highly available, multi-server environments.

  • Ability to troubleshoot infrastructure issues across operating system, networking, database and application layers.

  • Experience working alongside external infrastructure or managed service providers

Working Style & Soft Skills

  • Strong problem-solving and diagnostic capability.

  • Comfortable being given an outcome or problem rather than a predefined technical solution.

  • Able to independently investigate, evaluate options and implement appropriate solutions.

  • Strong ownership mentality and willingness to take responsibility for systems through their full lifecycle.

  • Proactive approach to identifying operational risks before they become incidents.

  • Calm and methodical approach to production incidents and high-pressure situations.

  • Strong attention to detail.

Good communication skills with both technical and non-technical stakeholders.

  • Ability to explain infrastructure risks and recommendations in practical business terms.

Nice-to-Have

  • Previous experience supporting SAP environments.

  • Experience working with infrastructure hosting or supporting SAP application and database workloads.

  • Experience with Terraform or similar Infrastructure as Code tooling.

  • Experience with cloud infrastructure.

Experience with centralised logging and observability platforms.

  • Experience with vulnerability management and security tooling.

  • Experience administering Redis or similar caching technologies.

  • Experience with highly available ecommerce or customer-facing infrastructure.

  • Experience supporting infrastructure across multiple physical locations or datacentres.

  • Experience working in organisations subject to formal audit, governance or listed-company control requirements.

  • Familiarity with broader Microsoft infrastructure and enterprise systems

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs4a.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all