Senior Platform Engineer

Cloud Gateway
London, UK
7 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Automation of Tests Bash Shell Ubuntu (Operating System) Cloud Computing Linux Monitoring of Systems Identity and Access Management Python (Programming Language) Linux System Administration
+27 more
Netconf Release Management Ansible Prometheus Systems Integration TypeScript Software Vulnerability Management Scripting Grafana Amazon Virtual Private Cloud (VPC) Juniper Database Migration Amazon Relational Database Service Influxdb Deployment Automation Bitbucket Fortinet Route53 Vertica Cloudwatch Api Gateway Software Coding Terraform Open Network Automation Platform Serverless Computing Docker Vulnerability Analysis

Job description

We’re looking for a Senior Platform Engineer to join our small but impactful platform team. You’ll

own and operate a multi-account AWS estate supporting critical government services, build and

maintain CI/CD pipelines, and contribute to our growing observability and network automation

products.

This is a hands-on role with real ownership. You won’t be pushing tickets around a queue; you’ll

be designing infrastructure, writing code, responding to incidents, and shipping improvements to

production. You’ll work closely with our NetOps, ServiceDesk, and engineering teams to keep

things running and make them better.

What You’ll Be Doing

  • Managing and evolving a multi-account AWS organisation with Terraform-managed resources across production and development environments.
  • Building, maintaining, and optimising CI/CD pipelines for infrastructure deployments, application releases, Docker image builds, and database migrations.
  • Operating and improving our observability stack (Grafana, InfluxDB, Prometheus, OpenTelemetry, CloudWatch) monitoring 200+ network devices and all AWS infrastructure.
  • Supporting and developing our Observability-as-a-Service product built on ClickHouse Cloud, HyperDX, and OpenTelemetry, including customer onboarding and tenant management.
  • Participating in on-call rotation, responding to production incidents, writing postmortem reports, and driving reliability improvements.
  • Contributing to security and compliance initiatives including Cyber Essentials Plus certification, IAM governance, and vulnerability management.
  • Building and maintaining golden AMI pipelines with automated patching, CIS hardening, and vulnerability scanning.
  • Supporting the Network Automation Platform initiative, integrating with FortiGate, Juniper, and NetBox for automated device configuration management.
  • Maintaining and supporting legacy infrastructure across established accounts, including investigating issues in systems, understanding how they were built, and planning gradual modernisation where appropriate., * The chance to own a platform end-to-end, not just a slice of it.
  • Real impact, the services you support help millions of people across the UK.
  • A small team where your contributions are visible and valued.
  • Exposure to a wide range of technologies across infrastructure, serverless, observability, and network automation.
  • A growing product engineering function with opportunities to shape the direction of new

Requirements

  • Strong hands-on experience with AWS (EC2, ECS, Lambda, RDS, VPC, IAM, CloudWatch, API Gateway, S3, Route53). This is an AWS-heavy environment and you’ll be working with these services daily.
  • Solid Terraform experience. We manage all infrastructure as code and you’ll be writing, reviewing, and maintaining Terraform across multiple accounts and projects.
  • Experience with CI/CD pipelines. We use Bitbucket Pipelines but the principles matter more than the specific tool. You should understand automated testing, quality gates, deployment strategies, and production approval workflows.
  • Docker experience. We run production container images on ECS with multi-stage builds and ECR.
  • Linux administration and troubleshooting. Amazon Linux and Ubuntu are our primary operating systems.
  • Experience with monitoring and observability tools. Grafana, Prometheus, InfluxDB, CloudWatch, or similar.
  • Scripting ability in at least one of TypeScript, Python, or Bash.
  • A security-conscious mindset. Our clients are government bodies and everything we build has to meet strict compliance standards.
  • Comfortable with on-call responsibilities and incident response.
  • Comfortable working with legacy infrastructure and services.
  • Experience with network automation tools (Ansible, NAPALM, Scrapli, NETCONF).
  • Familiarity with ClickHouse, OpenTelemetry, or time-series data platforms.
  • Experience working with government clients or in regulated environments.
  • Exposure to FinOps practices and cost optimisation.
  • AWS certifications (Solutions Architect Associate or similar).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all