Principal Data Engineer

Roche
Barcelona, Spain
8 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English, Spanish

Tech stack

Amazon Web Services Amazon S3 Software Quality Continuous Integration Data Architecture Information Engineering Data Governance Data Infrastructure Data Security Dimensional Modeling Python (Programming Language) SQL Databases
+6 more
Parquet Apache Spark IT Architecture Pyspark Terraform Data Pipelines

Job description

Experteer Overview As Principal Data Engineer at Roche, you will steer GenAI and advanced analytics strategy, shaping data architectures and scalable pipelines for diagnostic products.You lead technically while aligning with product and system architects to deliver high-impact, cost-efficient data initiatives.You will mentor teams, improve data governance and ensure compliant, secure data practices in a healthcare context.This role combines hands-on engineering with strategic leadership to advance Roche’s data-driven mission.Compensaciones / Beneficios * Define and drive the future strategy for GenAI and emerging technologies in diagnostic data analytics * Design, implement, and optimize data architectures using AWS services for performance and integration * Build and optimize data processing workflows with PySpark, SparkSQL, SQL, Iceberg, Parquet * Improve data infrastructure performance, scalability, and cost efficiency on AWS * Provide technical direction and mentorship, enforcing code quality and CI/CD * Collaborate with engineering and business stakeholders to translate requirements into data-driven solutions * Advocate for data security, governance, and regulatory compliance (HIPAA, GDPR) Responsabilidades * 5+ years of hands-on data engineering with leadership in technical projects * Expert-level AWS (S3, Redshift, Glue, Athena, EMR, Step Functions, MWAA) * Mastery of Python, SQL, and Spark * Experience with Dimensional Modeling, Data Partitioning, Terraform (IaC) * Ability to provide technical direction, architecture guidance, and cross-functional mentoring * Strong communication skills in English; Spanish is a plus * Bachelor’s or Master’s degree in Engineering Requisitos principales *

Requirements

Compensaciones / Beneficios * Define and drive the future strategy for GenAI and emerging technologies in diagnostic data analytics * Design, implement, and optimize data architectures using AWS services for performance and integration * Build and optimize data processing workflows with PySpark, SparkSQL, SQL, Iceberg, Parquet * Improve data infrastructure performance, scalability, and cost efficiency on AWS * Provide technical direction and mentorship, enforcing code quality and CI/CD * Collaborate with engineering and business stakeholders to translate requirements into data-driven solutions * Advocate for data security, governance, and regulatory compliance (HIPAA, GDPR) Responsabilidades * 5+ years of hands-on data engineering with leadership in technical projects * Expert-level AWS (S3, Redshift, Glue, Athena, EMR, Step Functions, MWAA) * Mastery of Python, SQL, and Spark * Experience with Dimensional Modeling, Data Partitioning, Terraform (IaC) * Ability to provide technical direction, architecture guidance, and cross-functional mentoring * Strong communication skills in English; Spanish is a plus * Bachelor’s or Master’s degree in Engineering Requisitos principales *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

Videos

See all

Related articles

See all