Senior Data Platform / Data Engineer

Straumann Group
Madrid, Spain
8 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Big Data Continuous Integration Data Infrastructure Data Security Python (Programming Language) PostgreSQL Machine Learning Meta-Data Management Raw Data Azure Machine Learning
+6 more
Software Engineering Management of Software Versions Data Ingestion Machine Learning Operations Data Lakehouse Data Pipelines

Job description

Straumann GroupLa siguiente información ofrece un resumen de las habilidades, cualidades y cualificaciones necesarias para este puesto.At Straumann Group we’re on an exciting journey of growth, innovation, and impact - driven by our mission to improve oral health and transform millions of lives worldwide.United by purpose, we bring our best selves to work every day, embracing a high-performance, player-learner culture that inspires collaboration, curiosity, and ambition.Here, you’ll have the opportunity to take charge of your own career, harnessing your skills, passion, and enthusiasm for learning to continually grow and progress.Together, we’re not just shaping brighter smiles, we’re unlocking the potential of people everywhere, including our own.About The RoleWe are looking for a Senior Data Platform / Data Engineer to join our ML Platform team and help build and scale the data infrastructure that powers our AI products in dentistry.Our platform supports the full AI development lifecycle, from raw data ingestion and annotation workflows to dataset versioning and model training pipelines.You will work closely with Machine Learning Researchers (MLRs), MLOps engineers, and product teams to ensure our data infrastructure is reliable, scalable, and easy to use.A key focus of the role is improving our Data Lakehouse (DLH) and dataset management workflows, including dataset versioning (DVC) and improving how data is prepared, extracted, and consumed across research and production systems.What You Will Work OnYou will play a key role in shaping the next generation of our data platform.Typical Responsibilities IncludeData platform ownershipDesign and evolve the Data Lakehouse (DLH) architecture used across our ML teams.Improve the reliability and structure of data ingestion, extraction, and transformation pipelines.Ensure datasets used for training and evaluation are consistent, reproducible, and well documented.Dataset lifecycle managementImprove workflows for dataset versioning and reproducibility using tools such as DVC.Design solutions for managing multiple versions of datasets and annotations across experiments and models.Improve the ability for researchers to retrieve the correct dataset versions reliably.Data pipelines and infrastructureBuild and maintain scalable data pipelines in Python.Improve metadata management, dataset validation, and data quality monitoring.Optimize data workflows across AWS-based infrastructure.Collaboration with ML teamsWork closely with ML researchers and ML engineers to understand their data needs.Support research workflows with reliable and efficient data access patterns.Help translate research requirements into robust platform capabilities.Data governance and qualityImplement practices for data quality, reproducibility, and traceability across the ML lifecycle.Ensure our data infrastructure meets the requirements of regulated AI development.xqbhyrx Must HaveWhat we’re looking for:Strong Python engineering skillsExperience building data pipelines or data platformsExperience working with AWSExperience working with large datasets used in ML workflowsStrong software engineering practices (testing, CI/CD, documentation)Experience collaborating with ML teams or working in AI environmentsNice To HaveExperience with dataset versioning tools such as DVCExperience with KubernetesExperience with data lakehouse architecturesExperience working with annotation pipelines or ML training datasetsExperience with PostgreSQL, Metabase, or similar data toolingExperience working in regulated environments (medical / healthcare AI)Our stackAWSPythonKubernetesPostgreSQLMetabaseDVC for dataset versioningInternal Data Lakehouse infrastructureAll qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or disability.Employment Type: Full TimeAlternative Locations: Spain : MadridTravel Percentage: 0 - 10%Requisition ID: **#J-**-Ljbffr

Requirements

datasets and annotations across experiments and models.Improve the ability for researchers to retrieve the correct dataset versions reliably.Data pipelines and infrastructureBuild and maintain scalable data pipelines in Python.Improve metadata management, dataset validation, and data quality monitoring.Optimize data workflows across AWS-based infrastructure.Collaboration with ML teamsWork closely with ML researchers and ML engineers to understand their data needs.Support research workflows with reliable and efficient data access patterns.Help translate research requirements into robust platform capabilities.Data governance and qualityImplement practices for data quality, reproducibility, and traceability across the ML lifecycle.Ensure our data infrastructure meets the requirements of regulated AI development. xqbhyrx Must HaveWhat we’re looking for:Strong Python engineering skillsExperience building data pipelines or data platformsExperience working with AWSExperience working with large datasets used in ML workflowsStrong software engineering practices (testing, CI/CD, documentation)Experience collaborating with ML teams or working in AI environmentsNice To HaveExperience with dataset versioning tools such as DVCExperience with KubernetesExperience with data lakehouse architecturesExperience working with annotation pipelines or ML training datasetsExperience with PostgreSQL, Metabase, or similar data toolingExperience working in regulated environments (medical / healthcare AI)Our stackAWSPythonKubernetesPostgreSQLMetabaseDVC for dataset versioningInternal Data Lakehouse infrastructureAll qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or disability.Employment Type: Full

About the company

Straumann GroupLa siguiente información ofrece un resumen de las habilidades, cualidades y cualificaciones necesarias para este puesto.At Straumann Group we’re on an exciting journey of growth, innovation, and impact - driven by our mission to improve oral health and transform millions of lives worldwide. United by purpose, we bring our best selves to work every day, embracing a high-performance, player-learner culture that inspires collaboration, curiosity, and ambition. Here, you’ll have the opportunity to take charge of your own career, harnessing your skills, passion, and enthusiasm for learning to continually grow and progress. Together, we’re not just shaping brighter smiles, we’re unlocking the potential of people everywhere, including our own.About The RoleWe are looking for a Senior Data Platform / Data Engineer to join our ML Platform team and help build and scale the data infrastructure that powers our AI products in dentistry. Our platform supports the full AI development lifecycle, from raw data ingestion and annotation workflows to dataset versioning and model training pipelines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:09 min

Challenges of interpreting raw data with language models

Clemens Vasters Clemens Vasters · World Congress 2025

5:37 min

Extensibility and programmability features of the PostgreSQL database

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:42 min

Building data pipelines and managing machine learning features

Hauke Brammer · World Congress 2021

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all