Infrastructure Engineer

Jobposting
London, UK
11 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Python (Programming Language) Prometheus Datadog Pulumi Google Cloud System Availability Large Language Models Grafana Kubernetes Infrastructure Automation Frameworks Information Technology
+4 more
Deployment Automation Terraform Stream Processing Data Pipelines

Job description

In this role, you will be the anchor for our European engineering cohort, drive system reliability, and serve as the infrastructure lead during EU business hours. You’ll balance active incident response and system maintenance with long-term infrastructure building., * Serve as the primary infrastructure contact during EU working hours, extending coverage to support cross-regional needs in APAC (AUS) and EMEA.

  • Build, scale, and optimize infrastructure dedicated to supporting our growing EMEA teams and regional customer deployments.
  • Proactively manage system health, monitoring, alerting, and incident response for GCP-based workloads to maintain high availability.
  • Automate deployments, manage Kubernetes clusters, and build internal tooling that empowers developers to ship code safely and fast.
  • Establish effective async workflows and handoff protocols with other infrastructure teams to ensure seamless 24/7 continuity.

Requirements

  • 5+ years of experience as a software engineer in a production environment.
  • Bachelor’s degree in computer science, engineering, or math.
  • 3+ years of hands-on experience running and scaling workloads in Google Cloud Platform (GCP).
  • 3+ years of experience managing production Kubernetes clusters (GKE experience is a big plus).
  • Proficiency in Python and Go for building infrastructure automation, APIs, and operational tooling.
  • Demonstrated experience with Infrastructure as Code (Terraform/Pulumi), CI/CD pipelines, and observability tools (Prometheus, Grafana, Datadog, etc.).
  • High degree of self-direction and clear async communication skills, with experience thriving in distributed team environments.
  • Familiarity with hosting, serving, or orchestrating agentic AI systems and LLM workflows.

Nice-to-Have

  • Experience building and scaling modern data pipelines or stream-processing architectures.
  • Experience working with early-stage technical companies

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · World Congress 2023

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all