IT & Service Delivery Engineer

Trakm8 Ltd
Coleshill, UK
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Job source

Tech stack

Proxmox Microsoft Windows Active Directory Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Application Performance Management JIRA Automation of Tests Bash Shell Border Gateway Protocol BitLocker Drive Encryption
+86 more
VoIP Business Systems Ubuntu (Operating System) Configuration Management Cyber Security Databases Continuous Integration Customer Data Management Data Centers Dynamic Host Configuration Protocol Linux Distributed File Systems Disaster Recovery Domain Name System (DNS) Multi-Factor Authentication Elasticsearch Perl (Programming Language) HAProxy Hypertext Transfer Protocols (HTTP) Icinga Subnetting Virtual Private Networks (VPN) Python (Programming Language) Kerberos (Protocol) Local Area Networks Lightweight Directory Access Protocols (LDAP) PostgreSQL Live Connect (Windows) Simple Mail Transfer Protocols Microsoft SQL Server Windows Servers MySQL Routing Nginx Public Key Infrastructure Windows PowerShell Procurement Software RabbitMQ Redis Remote Access Technology Ansible Prometheus Reverse Proxy Microsoft SharePoint Security Information and Event Management Software Deployment Transmission Control Protocol (TCP) Oracle Linux VMware VSphere Software Vulnerability Management Windows Desktop Data Logging Data Processing Scripting Load Balancing Delivery Pipeline Grafana Technical Debt Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Git Microsoft InTune Build Management Amazon Relational Database Service Containerization Selinux Kubernetes Information Technology Influxdb Cassandra Atlassian Tools Apache Kafka Bitbucket Data Management Route53 Graphite Vertica Api Gateway Puppet Veeam Software Version Control Wsus Atlassian Bamboo Docker Pagerduty Vmware

Job description

Based at Trakm8’s head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8’s telematics and optimisation platforms and of the corporate IT services. The role spans a mixed environment of corporate office, on-premises data centres and AWS. Customer platforms process data in real time, alongside the standard business IT services used by employees. Reliable delivery of both is critical., Service Availability & Incident Management

  • Monitor the health, capacity and performance of the telematics and optimisation platforms, corporate IT services and supporting infrastructure, taking action to maintain agreed levels of availability and throughput.
  • Investigate and resolve incidents across Linux and Windows servers, databases, networks, storage, container platforms, end user devices and application deployments, owning each from first report or alert through to root cause and permanent fix.
  • Carry out structured root cause analysis for recurring or significant issues, implementing corrective and preventative actions.
  • Communicate incident progress, risks and resolution clearly to technical and non-technical stakeholders.

Infrastructure Engineering & Delivery

  • Design, build, configure and maintain secure, resilient server infrastructure across data centre, virtualised and AWS environments.
  • Manage and improve containerised services, networking, storage, databases and load balancing.
  • Own backup, replication and restore testing, and contribute to capacity planning and disaster recovery, keeping recovery points and recovery times fit for purpose and evidenced.
  • Work with development, platform and service teams to support reliable application deployment and end-to-end service performance.

Windows, Directory & Endpoint Services

  • Administer Active Directory, Group Policy, DNS and internal certificate services across the domain.
  • Own server and workstation patching and endpoint protection coverage, including approvals, safeguard holds, deployment rings.
  • Administer accounts, access rights and licensing, and maintain inventory, imaging and software deployment tooling.
  • Respond directly to user requests and access issues, taking ownership through to resolution within the team’s operating model.
  • Maintain internal IT services including LAN, wireless, VOIP and mobile telephony, and support IT procurement, supplier management and hardware refresh.

Automation, Monitoring & Continuous Improvement

  • Develop and maintain automation and scripts for repeatable operational tasks.
  • Maintain effective monitoring, alerting and dashboards across metrics and log platforms, reviewing thresholds and coverage to identify issues early and reduce avoidable incidents.
  • Build self-healing and auto-remediation where it is safe to do so, for example automated rebalancing of workloads in response to queue lag or resource pressure.
  • Identify technical debt and resilience gaps and recommend proportionate improvements, making effective use of modern tooling including AI-assisted development and automation.

Security, Documentation & Collaboration

  • Operate infrastructure with a security-first approach, applying access controls, hardening and vulnerability remediation in line with company policies.
  • Support logging, monitoring and security tooling, and assist with investigation and remediation of security alerts.
  • Plan and deliver changes and project work through the RFC and release process, booking into agreed release windows and providing test, rollback and post-implementation evidence.
  • Produce recurring operational and service reporting covering availability, capacity, patching, backup and security, and maintain accurate asset, licence and configuration records.
  • Create and maintain clear technical documentation, runbooks and support procedures for new and existing solutions.

Technology Environment

The environment currently includes the following technologies. The list provides context for the role and is not intended to mean that experience in every technology is essential. The expectation is competence in several of these groups and the ability to pick up the rest.

Service delivery platform

  • Oracle Linux and Windows operating systems across VMware vSphere, Proxmox and AWS
  • Kubernetes and Docker, using MetalLB and Kong API gateway
  • MySQL, Cassandra (Scylla), CockroachDB, Redis, Kafka, Elasticsearch, PostgreSQL and RabbitMQ
  • Ansible and AWX, with scripting using Bash, Python or Perl
  • TCP/IP networking, DNS, routing, VPNs, BGP, SMTP, HTTP and HTTPS
  • Palo Alto firewalls; HAProxy, NGINX and dynamic DNS services
  • Grafana, Prometheus, Graphite (ClickHouse), InfluxDB and Icinga, with alerting into OpsGenie
  • ELK for application logging and Wazuh for platform host security monitoring
  • Git, Bitbucket, Jira and Bamboo CI/CD pipelines
  • iSCSI SAN, storage pools, volumes and snapshots
  • AWS services including EC2, Route53, RDS, ELB, VPC, Multi-AZ subnets, VPN, Transit Gateway and Customer Gateway

Corporate IT

  • Windows Server and Windows desktop estates, Active Directory, Group Policy, DNS, DHCP, DFS and file and print services with Linux (Ubuntu)
  • Microsoft 365 and Entra ID, Intune, including Exchange Online, SharePoint, Teams, multi-factor authentication and Conditional Access
  • VMware vSphere and vCenter, with Veeam Backup & Replication and object storage for offsite copies
  • Endpoint management, patching and software deployment tooling, with BitLocker disk encryption
  • CrowdStrike Falcon endpoint protection and UTMStack SIEM
  • Palo Alto firewalls managed through Panorama, with GlobalProtect remote access
  • LAN, wireless, VOIP and mobile telephony across multiple UK sites
  • Puppet for configuration management and PowerShell for scripting, with Icinga and OpsGenie alerting shared across both estates
  • Internal certificate services, TLS and PKI, split-horizon DNS, LDAP and Kerberos
  • MS SQL behind business systems, plus the Atlassian suite, service desk, intranet and digital signage platforms, * Endpoint management, software deployment and patching tooling such as PDQ, WSUS or Intune.
  • Backup and recovery tooling such as Veeam, including restore testing.
  • Providing IT support to end users directly, with the confidence to deal with colleagues at any level.

Automation and monitoring

  • Configuration management with Ansible, Puppet or AWX.
  • Monitoring, alerting and dashboarding with Icinga, Grafana, Prometheus or similar.
  • Infrastructure-as-code, CI/CD tooling or automated testing for infrastructure changes.
  • AI-assisted coding and automation tooling used to accelerate operational and scripting work.

Security and assurance

  • Firewall policy administration, ideally Palo Alto, and load balancing or reverse proxy work.
  • Endpoint protection, vulnerability management or SIEM platforms.
  • System build following NIST standards, hardening, SELinux and server firewalls.
  • Information security frameworks such as ISO 27001 or Cyber Essentials.

Ways of working

  • Working within a formal RFC and change window process, including rollback planning and evidence.
  • Moving live services off unsupported platforms and end-of-life operating systems without an outage.
  • Disaster recovery, capacity planning and service continuity testing., As a Trakm8 Group employee, you have the following H&S responsibilities:
  • To comply with the Group H&S policy and procedures
  • To take care of your own health and safety and that of people who may be affected by what you do (or what you do not do)
  • To co-operate with others on health and safety
  • To not interfere with, or misuse, anything provided for your health, safety or welfare
  • To wear/use the correct Protective Equipment at all times
  • To follow the training you have received at all times

Information Security:

As a Trakm8 Group employee, you have the following Information Security responsibilities:

  • To follow the Group Information Security policies at all times
  • To follow the Trakm8 Secure Development Policy at all times;
  • To ensure that passwords under your responsibility are kept confidential
  • To keep Group or Company information secure and confidential at all times
  • To inform your manager immediately if you detect, suspect or witness an incident that may be a breach of security

Requirements

  • Strong problem solving, with the ability to pick up an unfamiliar system quickly, retain it, and carry what you learn across into unrelated parts of the estate. Recognising that a fault in one area has the same shape as one you solved somewhere else entirely is worth more here than depth in any single technology.
  • Hands-on experience administering Linux in a production, business-critical environment, including fault diagnosis, performance management, patching and security hardening.
  • Solid Windows Server administration, including Active Directory and Group Policy.
  • Experience administering Microsoft 365 and Entra ID.
  • A working understanding of TCP/IP networking, DNS, routing, VPNs and firewall policy: enough to establish whether a fault sits in the application, the host, or a rule upstream.
  • A working understanding of Kubernetes and containerised services: enough to operate and diagnose them day to day. Deeper experience is welcome but not essential.
  • Scripting and automation in at least one of Bash, Python or PowerShell, and a preference for putting configuration into version control rather than doing things by hand.
  • Methodical, evidence-led troubleshooting: working from log, packet and configuration evidence rather than assumption, including faults that do not originate where they first appear.
  • Clear written and verbal communication, with the ability to explain technical issues and risks to different audiences.
  • A full UK driving licence is required, as occasional travel to data centres or other locations may be required to support IT and platform services.
  • Participation in the team’s 24x7 on-call rota, with flexible working outside normal hours where reasonably required to protect or restore critical services.

Desirable Skills / Experience:

Platform and cloud

  • Depth in Kubernetes administration, cluster operations and container networking.
  • AWS, VMware or other virtualisation and cloud platforms.
  • Databases in production, such as MySQL, PostgreSQL, MS SQL, Cassandra, Kafka or Elasticsearch.
  • API gateway technologies such as Kong and measuring or reporting API performance.
  • Telemetry, IoT, high-throughput data platforms or similarly critical real-time services.
  • Storage, SAN and data centre experience covering physical security, networking and compute., * Calm, methodical and solutions-focused, particularly when responding to live service issues.

  • Takes ownership and follows actions through to resolution.

  • Proactive in identifying risks, improvement opportunities and emerging capacity or resilience concerns.

  • Security-conscious and careful when working with critical systems and customer data.

  • Collaborative and approachable, with a willingness to share knowledge and support colleagues.

  • Organised and adaptable, able to balance planned work with changing operational priorities.

  • Committed to clear documentation, continuous learning and improving how the team works.

About the company

Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery. Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all