Principal Platform Engineer
Role details
Job location
Tech stack
Job description
An established and growing technology IT company near Cardiff is seeking a Principal Platform Engineer to take ownership of platform reliability, observability, automation, and day-to-day engineering coordination. This is a senior, hands-on role where you'll shape platform standards, improve operational efficiency, and act as the technical authority for critical platform services.
If you're passionate about platform reliability, automation, and leading engineering best practice, this role offers the opportunity to make a significant impact within a modern, forward-thinking environment., As Principal Platform Engineer, you will:
- Own the operational health, stability, and reliability of the platform estate
- Act as the senior technical escalation point for platform issues and major incidents
- Lead investigations into complex technical problems and drive service recovery
- Oversee observability tooling, monitoring standards, dashboards, and service health metrics
- Coordinate day-to-day workload, priorities, and activities of Platform Engineers
- Maintain engineering standards, documentation, and repeatable operational processes
- Drive automation to reduce manual effort and improve operational efficiency
- Support platform lifecycle activities including upgrades, migrations, optimisation and service improvements
- Contribute technical expertise to new platform capabilities and service development
Requirements
- Proven background in senior infrastructure, platform, or reliability engineering
- Strong hands-on experience with VMware, Hyper-V, and Veeam Backup & Replication
- Experience supporting business-critical platforms and major incident recovery
- Strong troubleshooting capability across infrastructure and platform services
- Experience with monitoring/observability tools (e.g., Grafana)
- Demonstrated ability to reduce manual effort through automation
- Experience coordinating technical workload within a small engineering team
- Understanding of backup, recovery, and resilience technologies
Desirable experience:
- Zerto, OpsRamp, Cohesity, Microsoft 365 backup solutions
- Workflow automation tools (e.g., n8n)
- Infrastructure as Code concepts
- MSP / Managed Services environment experience
- ITIL-aligned operational environments
- Supporting customer onboarding and service transition, * Customer-centric mindset
- Strong commitment and ownership
- Courage to challenge and improve existing processes
- Innovative, proactive, and solutions-focused
Ready to take the lead?