> Markdown version of [/jobs/ext/2790881-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2790881-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Engineer - **Company:** Shiftautomate - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, DevOps, Fault Tolerance, Python (Programming Language), Reliability Engineering, Data Logging, Pulumi, Kubernetes, Terraform - **Published:** September 8, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/infrastructure-engineer/46460395 ## About the Role design from conception through launchReport to the director of engineeringRequirements5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product companyExperience running containerisation in productionExperience with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)Good proficiency in Python or Go for automation and toolingDaily workflow already includes agentic tooling such as Claude Code, Droid, Codex, or internal skills; this is a hard requirementDemonstrated ability to challenge the status quo, identify systemic weaknesses, and propose innovative solutions to complex reliability problemsAbility to reason from constraints and failure modes and articulate tradeoffs in business termsAbility to make reversible decisions, write rollback plans, work with monitoring and logging stacks, and stress systems safelyExcellent communication, collaboration, and problem-solving skillsStrong ownership and accountability for mission-critical systemsAt least one end-to-end 0-to-1 infrastructure build with an attached outcome metricWillingness and ability to work in person in the office 3 days per weekLegally authorized to work in the country where the job is locatedCore CompetenciesDemonstrates expertise in building and operating resilient, scalable infrastructure for high-traffic enterprise platforms, with a strong focus on automation using Python or Go. Proven ability to lead incident response and uphold reliability standards while collaborating effectively with cross-functional teams.Highest-signal resume keywordsInfrastructure EngineeringDevOps PracticesAWS, GCP, AzurePython or Go AutomationHelm and TerraformHard SkillsInfrastructure EngineeringDevOpsAutomationContainerizationIncident ResponseRoot-Cause AnalysisSLO DefinitionMonitoring and LoggingSystem DesignHigh-Availability SystemsSoft SkillsExcellent CommunicationCollaborationProblem-SolvingOwnershipAccountabilityIndustry KeywordsGenerative AIHigh-Traffic PlatformsProduction SystemsReliability EngineeringHigh-Growth Product CompanyTools & TechnologiesKubernetesHelmTerraformPulumiCloud ToolingAgentic Tooling #J-18808-Ljbffr ## Description Build resilient, scalable, fault-tolerant infrastructure for WRITER's high-traffic enterprise generative AI platformMove between SRE, DevOps, Infrastructure, and Platform initiatives as priorities shiftAutomate operational tasks and infrastructure management with Python or GoDesign and operate infrastructure across AWS, GCP, and AzureWork with Kubernetes, Helm, Terraform, and cloud and AI toolingUse AI agents to investigate incidents, draft Terraform and Helm changes, write runbooks, scaffold tooling, and review pull requestsEncode recurring infrastructure tasks as reusable internal skills for human and agent teammatesLead incident response, post-mortems, and root-cause analysesOwn reliability, performance, and efficiency of core services end-to-endDefine and uphold SLOs and error budgets and carry the on-call pagerBalance immediate reliability work with long-term platform, observability, cost, and reliability investmentsCollaborate with product, security, and engineering peers on system ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)