Senior Site Reliability Engineer - AWS Kubernetes

Source Technology
London, UK
11 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Storage Continuous Integration DevOps Disaster Recovery File Systems Domain Name System (DNS) HAProxy Monitoring of Systems
+36 more
Hyper-V Networking Hardware Python (Programming Language) Network Security Network Troubleshooting Network Configuration and Change Management NetFlow Network Forensics Network Monitoring Routing Packet Analyzer Network Protocols Nginx Performance Tuning Reliability Engineering Prometheus Simple Network Management Protocols Wireshark Backup and Restore Scripting Google Cloud Load Balancing Cloud Monitoring Istio Grafana Firewalls (Computer Science) Cloudformation Containerization Kubernetes Linkerd (Service Mesh) Cloudwatch Terraform Splunk AWS EKS Docker Vmware

Job description

A truly unique opportunity to help launch a brand new team within a global financial services provider. This new team of highly skilled Full Stack Infrastructure Engineers will cover Compute, Storage, Network and Cloud technologies. You will help design, implement, and manage robust infrastructure solutions, ensuring reliability, scalability, and performance.

Requirements

  • Proven experience managing and optimizing a diverse infrastructure stack.
  • Extensive knowledge of cloud platforms (AWS, Azure, GCP) and infrastructure as code (Terraform, CloudFormation).
  • Familiarity of service mesh technologies (Istio, Linkerd).
  • Solid understanding of virtualization (VMware, Hyper-V) and containerization (Docker, Kubernetes) and orchestration.
  • Understanding of storage solutions (SAN, NAS, cloud storage) and backup systems.
  • Strong understanding of network protocols, routing, switching, and firewalls. * Experience with load balancers (F5, HAProxy, Nginx) and network monitoring tools.
  • Experience in DNS management and troubleshooting.
  • Experience in network security best practices.
  • Proficiency in monitoring and observability tools (Prometheus, Grafana, Splunk).
  • Proficiency in at least one scripting language (Python, Bash) for automation.
  • Experience with CI/CD pipeline management and DevOps practices.
  • Strong understanding of disaster recovery and business continuity planning.
  • Experience with performance tuning and capacity planning.
  • Understanding of chaos engineering principles and practices.
  • Skills in cost optimization for cloud infrastructure.

Specific Tools and Techniques:

  • Experience in using cloud native monitoring tools like AWS CloudWatch, Azure Monitor, and Google Cloud Operations Suite.
  • Experience with packet capture tools like Wireshark for troubleshooting network issues.
  • Experience in using traceroute utilities and performance analysis tools like perf for identifying and resolving bottlenecks.
  • Familiarity with tools such as ipconfig/ifconfig for viewing network configurations, flushing DNS, and diagnosing network issues.
  • Experience with SNMP-based tools for network device monitoring and performance management.
  • Experience in using NetFlow for network traffic analysis.
  • Experience with tools like iostat, vmstat, and dstat for monitoring storage and system performance.
  • Experience in tools like df, du, lsblk, and fdisk for managing and troubleshooting file systems and disk partitions.
  • Familiarity with tools like Prometheus and Grafana for monitoring and observability

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:39 min

Modernizing operational excellence and cloud deployment tools

Mustafa Toroman · World Congress 2023

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

Videos

See all

Related articles

See all