> Markdown version of [/jobs/ext/3242790-senior-infrastructure-engineer-i](https://www.wearedevelopers.com/jobs/ext/3242790-senior-infrastructure-engineer-i). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Infrastructure Engineer I - **Company:** Thought Machine - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Databases, Disaster Recovery, Fault Tolerance, Python (Programming Language), PostgreSQL, Open Source Technology, Prometheus, Software Engineering, Data Streaming, Systems Integration, Google Cloud, Grafana, Reliability of Systems, Kubernetes, Information Technology, Apache Kafka, Docker, Golang - **Published:** September 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=50f3f0316203c1b3 ## About the Role * Degree in Computer Science, Engineering, or a similar technical field. * 5+ years of hands-on software engineering experience building and scaling infrastructure platforms (not just maintaining them). * Strong proficiency in building production-ready platform tooling using Python or Golang. * Deep understanding of container and orchestration internals (e.g., Kubernetes, Docker). * Hands-on experience extending and integrating open-source tools like Prometheus, OpenTelemetry, and Grafana. * Proven ability to package and scale developer-facing infrastructure utilities with a focus on improving system reliability. Desirable * Experience building cloud-agnostic solutions running across AWS, Azure, and GCP. * Technical depth in databases or event-streaming platforms (e.g., Kafka, Postgres, DuckDB). * A deep interest in the internals of infrastructure tools, with a track record of troubleshooting, extending, or optimizing them within a professional environment. ## Description At Thought Machine, our Senior Software Engineers specialised in Infrastructure play a crucial role in deploying and maintaining our cutting-edge software, developing tools and integrations with third-party infrastructure components used in our customers' environments. Our team's mission is to abstract away multi-cloud complexity and reduce cognitive load for other developers. We treat our platform infrastructure which spans, orchestration, observability, and data. As a first-class product, enabling our internal teams and global clients to operate securely and reliably at massive scale. Duties * Scale the Core Platform: Evolve, scale, and optimise the cloud-agnostic platform infrastructure powering Vault globally. * Simplify Orchestration: Build intuitive control planes and automation to further abstract away infrastructure complexity, significantly easing the cognitive load for Thought Machine engineers and bank SREs. * Design for Resiliency: Engineer robust, fault-tolerant architectures and Disaster Recovery / Business Continuity Planning (DR/BCP) systems. * Advance Observability: Enhance our comprehensive observability suites that are shared across the entire Thought Machine product ecosystem, providing clients with deep, out-of-the-box monitoring and insights * Optimise Data & Streaming: Scale and optimise core databases and event-streaming platforms to achieve maximum performance and scalability at the lowest cost ## Related Videos - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)