Infrastructure Engineer

Shiftautomate
London, UK
12 days ago
Apply on www.apply4u.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing DevOps Fault Tolerance Python (Programming Language) Reliability Engineering Data Logging Pulumi Kubernetes Terraform

Job description

Build resilient, scalable, fault-tolerant infrastructure for WRITER’s high-traffic enterprise generative AI platformMove between SRE, DevOps, Infrastructure, and Platform initiatives as priorities shiftAutomate operational tasks and infrastructure management with Python or GoDesign and operate infrastructure across AWS, GCP, and AzureWork with Kubernetes, Helm, Terraform, and cloud and AI toolingUse AI agents to investigate incidents, draft Terraform and Helm changes, write runbooks, scaffold tooling, and review pull requestsEncode recurring infrastructure tasks as reusable internal skills for human and agent teammatesLead incident response, post-mortems, and root-cause analysesOwn reliability, performance, and efficiency of core services end-to-endDefine and uphold SLOs and error budgets and carry the on-call pagerBalance immediate reliability work with long-term platform, observability, cost, and reliability investmentsCollaborate with product, security, and engineering peers on system

Requirements

design from conception through launchReport to the director of engineeringRequirements5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product companyExperience running containerisation in productionExperience with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)Good proficiency in Python or Go for automation and toolingDaily workflow already includes agentic tooling such as Claude Code, Droid, Codex, or internal skills; this is a hard requirementDemonstrated ability to challenge the status quo, identify systemic weaknesses, and propose innovative solutions to complex reliability problemsAbility to reason from constraints and failure modes and articulate tradeoffs in business termsAbility to make reversible decisions, write rollback plans, work with monitoring and logging stacks, and stress systems safelyExcellent communication, collaboration, and problem-solving skillsStrong ownership and accountability for mission-critical systemsAt least one end-to-end 0-to-1 infrastructure build with an attached outcome metricWillingness and ability to work in person in the office 3 days per weekLegally authorized to work in the country where the job is locatedCore CompetenciesDemonstrates expertise in building and operating resilient, scalable infrastructure for high-traffic enterprise platforms, with a strong focus on automation using Python or Go. Proven ability to lead incident response and uphold reliability standards while collaborating effectively with cross-functional teams.Highest-signal resume keywordsInfrastructure EngineeringDevOps PracticesAWS, GCP, AzurePython or Go AutomationHelm and TerraformHard SkillsInfrastructure EngineeringDevOpsAutomationContainerizationIncident ResponseRoot-Cause AnalysisSLO DefinitionMonitoring and LoggingSystem DesignHigh-Availability SystemsSoft SkillsExcellent CommunicationCollaborationProblem-SolvingOwnershipAccountabilityIndustry KeywordsGenerative AIHigh-Traffic PlatformsProduction SystemsReliability EngineeringHigh-Growth Product CompanyTools & TechnologiesKubernetesHelmTerraformPulumiCloud ToolingAgentic Tooling #J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · World Congress 2023

Videos

See all

Related articles

See all