NOC Engineer, AWS

Spectrum IT Recruitment
Reading, UK
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Job source

Tech stack

Amazon Web Services Bash Shell Cloud Computing Domain Name System (DNS) Python (Programming Language) Linux System Administration Networking Basics Prometheus TCP/IP Datadog Load Balancing Grafana
+5 more
Kubernetes Cloudwatch Terraform Splunk Docker

Job description

Monitoring and maintaining highly available production platforms running in AWS Responding to and managing production incidents across a 24/7 service Investigating complex technical issues and restoring services quickly and effectively Developing automation to reduce manual operational tasks and improve platform resilience Building and improving monitoring, alerting and observability across cloud environments Working alongside Software, Platform, Cloud and Security Engineers to improve reliability and operational excellence Contributing to post-incident reviews and driving continuous service improvements Supporting containerised workloads using Kubernetes and Docker

Requirements

You’ll ideally have experience in a Production Engineering, Cloud Operations or NOC environment with exposure to:

Linux systems administration AWS cloud infrastructure Kubernetes and Docker Production support and incident management Python, Bash or Go scripting Monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch Networking fundamentals including DNS, TCP/IP and load balancing A passion for automation, continuous improvement and operational excellenceExperience with Infrastructure as Code (Terraform), SRE principles (SLIs, SLOs), or regulated environments would be beneficial but isn’t essential.

Benefits & conditions

Fully Remote (UK) 24/7 Shift Pattern (28-day rota including days & nights) £ Competitive + Bonus + Excellent Benefits

Build resilient cloud platforms that support critical national services.

This is far more than a traditional NOC role.

You’ll be joining an engineering-led organisation where reliability, automation and continuous improvement sit at the heart of the platform. Rather than simply responding to incidents, you’ll work to prevent them by improving systems, automating operational processes and helping shape the future of highly resilient cloud services.

If you’re passionate about building reliable cloud platforms and enjoy solving complex technical problems in large-scale production environments, we’d love to hear from you.

What you’ll be doing

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on itjobpro.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:23 min

Reviewing AWS infrastructure deployment configuration and planning

Devlin Duldulao · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all