Site Reliability Engineer

LexisNexis Risk Solutions
Carshalton, UK
7 days ago
Apply on www.careerjet.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Distributed Systems Domain Name System (DNS) Github Identity and Access Management Python (Programming Language) Key Management Networking Basics OpenID
+15 more
Reliability Engineering Azure Active Directory Cloud Services Prometheus Lexis Datadog Data Logging Load Balancing Cloud Monitoring Delivery Pipeline Grafana Kubernetes Cloudwatch Terraform Golang

Job description

Are you passionate about building reliable, scalable cloud platforms that empower engineering teams to deliver at speed? Do you enjoy solving complex infrastructure challenges, driving automation, and improving operational excellence across a modern cloud environment? About the Business LexisNexis® Risk Solutions provides customers with solutions and decision tools that combine public and industry specific content with advanced technology and analytics to assist them in evaluating and predicting risk and enhancing operational efficiency. We use the power of data and advanced analytics to help our customers make better, timelier decisions. By bringing clarity to information, we ultimately help make communities safer, insurance rates more accurate, commerce more transparent, business decisions easier and processes more efficient. You can learn more about LexisNexis Risk at https://risk.lexisnexis.com/ About our Team Our team is responsible for delivering complex, highly opinionated infrastructure as code capabilities that underpins how applications are deployed across our organisation spanning different business units and use cases. Terraform modules and delivery pipelines to ensure teams can provision feature rich Kubernetes clusters that are ready to operate workloads in a safe and consistent way., We’re looking for a Site Reliability Engineer to design, build and operate cloud infrastructure across AWS and Azure. You’ll play a key role in enabling engineering teams through infrastructure as code, automation, and platform reliability practices. This is a hands-on opportunity to help shape scalable, secure, and observable cloud platforms while driving engineering excellence across teams. Responsibilities

  • Design, implement and maintain infrastructure as code for use across cloud environments
  • Manage and operate Kubernetes clusters (EKS/AKS)
  • Build and maintain CI/CD pipelines to automate build, test and deployment workflows
  • Own the reliability, security and cost-efficiency of cloud infrastructure
  • Troubleshoot and resolve infrastructure issues and fix root causes of issues at the source
  • Implement monitoring, logging and alerting solutions to ensure system observability
  • Collaborate with engineers and architects to improve platform documentation, standards and adoption

Requirements

  • Strong hands-on experience with cloud platforms (AWS & Azure ideally)
  • Hands-on experience with Infrastructure as Code in production environments using Terraform or similar
  • Be able to diagnose and resolve complex infrastructure issues across distributed systems
  • Have experience building and maintaining CI/CD pipelines using GitHub Actions or similar
  • Be comfortable writing automation scripts using Go, Python, Bash or similar
  • Understanding of networking fundamentals (VPCs/VNets, load balancers, DNS, security groups/NSGs)
  • Experience with secrets management and identity/access control (IAM, OIDC, Azure AD)
  • Familiarity with observability tooling (Prometheus, Grafana, CloudWatch, Azure Monitor)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.co.uk
Prepare application

Good distractions

Loading talks and stories from around this role…