Principal Platform Engineer

100% IT RECRUITMENT Limited
Cardiff, UK
24 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£60,000.0
Working hours
Regular working hours

Tech stack

Microsoft Windows Backup Devices Hyper-V Reliability Engineering Service Development Studio Datadog Grafana Nintex Veeam Vmware

Job description

An established and growing technology IT company near Cardiff is seeking a Principal Platform Engineer to take ownership of platform reliability, observability, automation, and day-to-day engineering coordination. This is a senior, hands-on role where you’ll shape platform standards, improve operational efficiency, and act as the technical authority for critical platform services.

If you’re passionate about platform reliability, automation, and leading engineering best practice, this role offers the opportunity to make a significant impact within a modern, forward-thinking environment., As Principal Platform Engineer, you will:

  • Own the operational health, stability, and reliability of the platform estate
  • Act as the senior technical escalation point for platform issues and major incidents
  • Lead investigations into complex technical problems and drive service recovery
  • Oversee observability tooling, monitoring standards, dashboards, and service health metrics
  • Coordinate day-to-day workload, priorities, and activities of Platform Engineers
  • Maintain engineering standards, documentation, and repeatable operational processes
  • Drive automation to reduce manual effort and improve operational efficiency
  • Support platform lifecycle activities including upgrades, migrations, optimisation and service improvements
  • Contribute technical expertise to new platform capabilities and service development

Requirements

  • Proven background in senior infrastructure, platform, or reliability engineering
  • Strong hands-on experience with VMware, Hyper-V, and Veeam Backup & Replication
  • Experience supporting business-critical platforms and major incident recovery
  • Strong troubleshooting capability across infrastructure and platform services
  • Experience with monitoring/observability tools (e.g., Grafana)
  • Demonstrated ability to reduce manual effort through automation
  • Experience coordinating technical workload within a small engineering team
  • Understanding of backup, recovery, and resilience technologies

Desirable experience:

  • Zerto, OpsRamp, Cohesity, Microsoft 365 backup solutions
  • Workflow automation tools (e.g., n8n)
  • Infrastructure as Code concepts
  • MSP / Managed Services environment experience
  • ITIL-aligned operational environments
  • Supporting customer onboarding and service transition, * Customer-centric mindset
  • Strong commitment and ownership
  • Courage to challenge and improve existing processes
  • Innovative, proactive, and solutions-focused

Ready to take the lead?

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · WWC 2021

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · WWC 2024

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:46 min

The history of abstractions and hardware virtualization

Edoardo Dusi · LIVE

Videos

See all

Related articles

See all