> Markdown version of [/jobs/ext/1476248-data-engineer](https://www.wearedevelopers.com/jobs/ext/1476248-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** LiveEO GmbH - **Location:** Berlin, Germany - **Contract:** Permanent contract - **Skills:** Geographic Information Systems, Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Cloud Computing, Information Engineering, Identity and Access Management, Python (Programming Language), PostgreSQL, Metadata, PostGIS, Spatial Data Infrastructures, SQLAlchemy, Management of Software Versions, Data Logging, Backend, Fastapi, Geospatial Data Abstraction Library (GDAL), Machine Learning Operations, Functional Programming - **Published:** July 29, 2026 - **Apply:** https://www.adzuna.de/details/5818281882 ## About the Role * Strong Python in production, not scripting-only. * Building and running backend services with FastAPI (or similar), SQLAlchemy, Alembic and Pydantic. * PostgreSQL, and spatial data (PostGIS, GeoAlchemy2 or equivalent) is a strong plus. * AWS in production, mainly S3, Lambda, EC2 and IAM basics. * Orchestration with Prefect, Airflow, Dagster or equivalent. * Comfort handling geospatial raster and vector data (rasterio, geopandas, shapely, STAC, COG) at a level where you can move it reliably. * Pragmatic delivery and a reliability and ops mindset, so you ship robust, well-tested services and keep them running. * Ownership and verification mindset, so you validate outputs against reality and catch results that look right but are not, using judgment and not just code. * Distributed compute with Ray or Anyscale is a plus. * MLflow, or experiment and dataset versioning is a plus. * Remote-sensing or geospatial foundations, or SAR exposure (Capella, Sentinel-1) is a plus. * Observability with structured logging, OpenTelemetry and alerting is a plus. * Satellite provider APIs such as Capella, UP42, Planet and ICEYE is a plus. ## Description We are looking for a Data Engineer (f/m/x) to own and harden the data-delivery and serving backbone of SurfaceScout, LiveEO's solution for monitoring changes around critical assets like pipelines and power grids to protect utility customers from third-party interference. You will make the path from detection to delivered customer insight reliable and repeatable. This covers the delivery-management and insight-generation services, the post-processing steps, and the AWS and Prefect platform that runs them end to end, at scale across thousands of km at a daily cadence. While our remote-sensing and detection-model specialists focus on imagery and models, you will be the custodian of the systems that turn their output into a dependable product that is deterministic, observable and resilient. This is a hands-on, delivery-first backend and platform role, with plenty of room to grow into the wider geospatial data platform. You will sit within the SurfaceScout Data Engineering & ML squad and work alongside the senior engineers who own the delivery and serving platform, sharing that ownership so the systems are no longer a single-person dependency. You will collaborate with the App squad and the platform teams that co-own the shared Imagery API, and partner with our in-house labelling studio, Operations Support Studio (OST), on data quality. SAR change-detection research, meaning coherent change detection, InSAR and semantic-change work, is driven by our remote-sensing specialist together with LiveEO's Section 4 team. You consume its outputs, you do not build them. Tech stack & tools: * Core language: Python (uv-managed). * Backend services: FastAPI, Uvicorn, SQLAlchemy, Alembic and Pydantic. * Data: PostgreSQL and PostGIS (via GeoAlchemy2). * Cloud and infrastructure: AWS, mainly S3, Lambda and EC2. * Orchestration and compute: Prefect and Ray/Anyscale. * Geospatial and EO: GDAL, Rasterio, GeoPandas, Shapely, STAC (pystac) and Cloud-Optimised GeoTIFF (rio-cogeo). * ML lifecycle: MLflow. Your challenge Delivery & serving services * Co-own and extend the delivery-management service (SDMS) and the insight-generation and post-processing services, the path that turns model detections into delivered, customer-ready insights. * Harden these into well-tested, documented services so they are no longer a single-person dependency. Platform, orchestration & AWS * Own and maintain the AWS and Prefect delivery platform, meaning the Lambda-based glue, the orchestration flows and the cross-repo release, and keep it deterministic and repeatable. * Manage the PostgreSQL and PostGIS data model behind delivery, including the schema and the migrations via Alembic. Reliability & delivery operations * Be a reliable second owner for where the pipeline breaks. You unblock stuck orchestration runs, keep delivery SLAs met, and cut out manual steps. Data quality, metadata & diagnostics * Automate QA across the serving path, covering schema and geometry integrity and coverage gaps, and maintain structured metadata and STAC entries so every delivery is traceable. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Building and Deploying Multi-Agent Systems with ADK and Vertex AI](https://www.wearedevelopers.com/videos/1918-building-and-deploying-multi-agent-systems-with-adk-and-vertex-ai) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)