Data Engineer
Role details
Job location
Tech stack
Job description
We are looking for a Data Engineer (f/m/x) to own and harden the data-delivery and serving backbone of SurfaceScout, LiveEO's solution for monitoring changes around critical assets like pipelines and power grids to protect utility customers from third-party interference. You will make the path from detection to delivered customer insight reliable and repeatable. This covers the delivery-management and insight-generation services, the post-processing steps, and the AWS and Prefect platform that runs them end to end, at scale across thousands of km at a daily cadence. While our remote-sensing and detection-model specialists focus on imagery and models, you will be the custodian of the systems that turn their output into a dependable product that is deterministic, observable and resilient. This is a hands-on, delivery-first backend and platform role, with plenty of room to grow into the wider geospatial data platform.
You will sit within the SurfaceScout Data Engineering & ML squad and work alongside the senior engineers who own the delivery and serving platform, sharing that ownership so the systems are no longer a single-person dependency. You will collaborate with the App squad and the platform teams that co-own the shared Imagery API, and partner with our in-house labelling studio, Operations Support Studio (OST), on data quality. SAR change-detection research, meaning coherent change detection, InSAR and semantic-change work, is driven by our remote-sensing specialist together with LiveEO's Section 4 team. You consume its outputs, you do not build them.
Tech stack & tools:
- Core language: Python (uv-managed).
- Backend services: FastAPI, Uvicorn, SQLAlchemy, Alembic and Pydantic.
- Data: PostgreSQL and PostGIS (via GeoAlchemy2).
- Cloud and infrastructure: AWS, mainly S3, Lambda and EC2.
- Orchestration and compute: Prefect and Ray/Anyscale.
- Geospatial and EO: GDAL, Rasterio, GeoPandas, Shapely, STAC (pystac) and Cloud-Optimised GeoTIFF (rio-cogeo).
- ML lifecycle: MLflow.
Your challenge
Delivery & serving services
- Co-own and extend the delivery-management service (SDMS) and the insight-generation and post-processing services, the path that turns model detections into delivered, customer-ready insights.
- Harden these into well-tested, documented services so they are no longer a single-person dependency.
Platform, orchestration & AWS
- Own and maintain the AWS and Prefect delivery platform, meaning the Lambda-based glue, the orchestration flows and the cross-repo release, and keep it deterministic and repeatable.
- Manage the PostgreSQL and PostGIS data model behind delivery, including the schema and the migrations via Alembic.
Reliability & delivery operations
- Be a reliable second owner for where the pipeline breaks. You unblock stuck orchestration runs, keep delivery SLAs met, and cut out manual steps.
Data quality, metadata & diagnostics
- Automate QA across the serving path, covering schema and geometry integrity and coverage gaps, and maintain structured metadata and STAC entries so every delivery is traceable.
Requirements
- Strong Python in production, not scripting-only.
- Building and running backend services with FastAPI (or similar), SQLAlchemy, Alembic and Pydantic.
- PostgreSQL, and spatial data (PostGIS, GeoAlchemy2 or equivalent) is a strong plus.
- AWS in production, mainly S3, Lambda, EC2 and IAM basics.
- Orchestration with Prefect, Airflow, Dagster or equivalent.
- Comfort handling geospatial raster and vector data (rasterio, geopandas, shapely, STAC, COG) at a level where you can move it reliably.
- Pragmatic delivery and a reliability and ops mindset, so you ship robust, well-tested services and keep them running.
- Ownership and verification mindset, so you validate outputs against reality and catch results that look right but are not, using judgment and not just code.
- Distributed compute with Ray or Anyscale is a plus.
- MLflow, or experiment and dataset versioning is a plus.
- Remote-sensing or geospatial foundations, or SAR exposure (Capella, Sentinel-1) is a plus.
- Observability with structured logging, OpenTelemetry and alerting is a plus.
- Satellite provider APIs such as Capella, UP42, Planet and ICEYE is a plus.
Benefits & conditions
- The opportunity to create a product that can improve business processes and lives across the globe.
- Flexible working hours and hybrid work model - we trust our employees to get their work done while maintaining a healthy work-life balance.
- We empower employees to drive their own career development, take initiative and have the freedom to be creative and bold.
- Not an overtime culture - we take care that overtime is done only as a necessity and always offset with time off and rest.
- A collaborative and learning environment - frequent internal workshops, knowledge sharing sessions, journal clubs and hackathons.
- Office located in the centre of Berlin Kreuzberg with free fruit, nuts and drinks.
- Potential to participate in the employee stock option program.
- Urban Sports membership and BVG subsidy, corporate pension program.
- A diverse and vibrant international environment of 30+ different nationalities