> Markdown version of [/jobs/ext/2019895-senior-data-engineer-data-platform-ontology](https://www.wearedevelopers.com/jobs/ext/2019895-senior-data-engineer-data-platform-ontology). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # (Senior) Data Engineer - Data Platform & Ontology - **Company:** Statista GmbH - **Location:** Berlin, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Data Analysis, Apache HTTP Server, Automation of Tests, Hospital Information Systems, Cloud Computing, Cloud Database, Databases, Continuous Integration, Information Engineering, Data Governance, Data Infrastructure, Github, Graph Database, JSON, Python (Programming Language), Metadata, Neo4j, Simple Data Format, Software Engineering, SQL Databases, Data Streaming, Data Storage Technologies, Cloud Platform System, Data Ingestion, Fast Healthcare Interoperability Resources, Delivery Pipeline, Snowflake, Reliability of Systems, Infrastructure as Code (IaC), Integration Tests, Information Technology, Data Lineage, Data Management, Restful APIs, Terraform, Data Pipelines - **Published:** August 11, 2026 - **Apply:** https://www.xing.com/jobs/berlin-senior-data-engineer-data-platform-ontology-157069857 ## About the Role Core Requirements (Must-Haves) * Data Ingestion & Pipeline Orchestration: Advanced Python and analytical SQL for complex data ingestion across diverse file formats, REST APIs, databases, and cloud lakes (S3/Iceberg). Hands-on experience with modern orchestrators (Prefect, Airflow, or Dagster). * Entity Resolution & Data Governance: Practical experience with entity resolution/record linkage frameworks (e.g., Splink, dedupe, recordlinkage) and schema management/data contracts (Pydantic, dbt contracts, or JSON Schema). * Cloud Platform & Warehouse Infrastructure: Deep hands-on experience in an AWS production environment (S3, ECS/EC2) combined with cloud data warehouses (Snowflake). * Automated CI/CD & Workflow Automation: Proven track record of automating data pipeline deployments, integration tests, and validation workflows via GitHub Actions. Nice-to-Haves (What Will Make You Stand Out) * Knowledge Graphs & Healthcare Terminologies: Exposure to ontology/semantic frameworks (RDF/OWL, SKOS, Neo4j, LinkML) or international medical classifications/vocabularies (SNOMED CT, ICD/OPS, FHIR). * Metadata & Lineage Tooling: Experience operating metadata registries and lineage catalogs (e.g., OpenMetadata, DataHub, dbt docs). * Infrastructure as Code (IaC): Proficiency in using Terraform to declaratively manage cloud resources and environments., * Degree: Bachelor's or Master's in Computer Science, Data Science, Software Engineering, or a related quantitative field. * Experience: 5+ years in data engineering building production pipelines and data platforms; including a sustained period within one organization seeing a core platform or product through build * launch * iteration. * Domain Knowledge: Healthcare domain experience is a plus (basic understanding of healthcare KPIs, quality metrics, or benchmarking concepts; familiarity with hospital structures and medical classification systems like ICD/OPS is especially valuable). * Mindset: Strong analytical and systems mindset, with a proven ability to transform messy, heterogeneous international data into a clean, well-governed, and highly structured data asset. * Languages: Fluent in English, German is a plus. * Working Style: Highly structured, curious, detail-oriented, and motivated to collaborate closely with analytics engineers, data scientists, and methodology experts in an international environment. ## Description As the Senior Data Engineer on our Healthcare Platform, you will own the foundational data ingestion, entity-resolution, and platform infrastructure end-to-end. You will design and operate scalable batch/streaming data pipelines, harden platform orchestration, and ensure system reliability, security, and cost efficiency. A central challenge of this role is building and operationalizing our unified healthcare data ontology-a consistent semantic model covering hospitals, departments, specialties, metrics, and classification standards. Working closely with Analytics Engineers, Data Scientists, and Methodology experts, you will turn heterogeneous, multi-country hospital data into a coherent, highly queryable data asset., * Pipeline Infrastructure & Orchestration: Build, optimize, and operate reliable ELT pipelines (using Python, SQL, and Prefect/Airflow) to ingest data from heterogeneous international sources, APIs, databases, and lakehouse storage (S3, Apache Iceberg). * Healthcare Ontology & Entity Resolution: Drive the implementation of entity resolution and master data management (MDM) for international hospital entities, mapping raw source data to canonical structures and maintaining standardized vocabularies (e.g., ICD/OPS, specialty taxonomies). * Data Contracts & Schema Governance: Establish strict data contracts (Pydantic, dbt contracts) and schema management, ensuring dataset reproducibility, data lineage tracking, and automated validation across all platform pipelines. * Platform Efficiency & Cloud Infrastructure: Optimize data storage, query execution, and compute costs across AWS and Snowflake, keeping data assets performant, secure, and cost-effective. * Automation & CI/CD: Implement automated testing and deployment workflows for data pipelines using GitHub Actions and Infrastructure as Code (Terraform). * Cross-Functional Data Enablement: Partner directly with Analytics Engineers, Data Scientists, and domain experts to deliver documented, research-grade, and production-ready datasets. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Data Analyst Salary Germany](https://www.wearedevelopers.com/magazine/277-data-analyst-salary-germany) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [Fullstack developer salary in Germany [2023]](https://www.wearedevelopers.com/magazine/197-fullstack-developer-salary-in-germany-2023) - [Backend Developer Salary in Germany [2023]](https://www.wearedevelopers.com/magazine/196-backend-developer-salary-in-germany-2023) - [Software Developer Salary in Germany [2023]](https://www.wearedevelopers.com/magazine/194-software-developer-salary-in-germany-2023)