> Markdown version of [/jobs/ext/1678284-data-engineer](https://www.wearedevelopers.com/jobs/ext/1678284-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Veeva Link - **Location:** Barcelona, Spain - **Salary:** €65,000.0 - €110,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Architectural Patterns, Cloud Engineering, Information Engineering, Data Integrity, Data Structures, Python (Programming Language), Machine Learning, Data Logging, Large Language Models, Apache Spark, Data Lakes, Pyspark, Solid Principles, Veeva, Data Pipelines - **Published:** July 13, 2026 - **Apply:** https://www.jobleads.com/es/job/e681627ddd578536711b89c7dfaaa04a1 ## About the Role * Experience with Python and Apache Spark/PySpark to handle massive datasets * Expertise in building cloud-native software within AWS or GCP * Background in designing and maintaining modern architectures, specifically Data Lakes, lakehouses and warehouses (DeltaLake, Redshift) * Experience operating LLM systems in production, including third-party model providers, human/data feedback loops, and multi-model traffic orchestration * Driving technical execution within Agile environments, utilizing strong English communication skills to align with global stakeholders ## Description At Veeva Link, we are building the intelligence layer for life sciences, creating connected data applications that accelerate drug development and significantly improve patient outcomes. Our core belief is that combining the highest quality data with state-of-the-art software delivers immense value. As a Data Engineer, you will be responsible for the life cycle of the data that defines the healthcare landscape. You will design, build, and maintain the robust data pipelines required to ingest and process global Healthcare Organization (HCO) data. You will be a key architect in managing the complex hierarchical relationships of over 4 million entities, ensuring data integrity, scalability, and seamless delivery to downstream stakeholders. We are an AI-forward team and actively promote AI-driven development practices, leveraging LLMs and automation to accelerate coding, optimize pipelines, and stay at the forefront of the evolving data landscape. * You will architect the "Data DNA" used by global biopharmas to make data-driven decisions * You will be at the forefront of enabling global AI initiatives within high-stakes, high-impact environments * You will design PySpark pipelines and collaborate on ML models to integrate diverse data into a robust lakehouse architecture * Identify, implement, and maintain end-to-end HCO data pipelines * Refine data structures and processing logic to meet the rapidly changing demands of the market * Deploy solid principles and clean patterns to data engineering tasks * Advance the long-term architectural roadmap * Validate high quality and availability of HCO deliveries * Govern and optimize the underlying infrastructure * Own monitoring, logging, and performance metrics from day one * A proactive interest in using AI tools to streamline development and solve complex data problems ## Related Videos - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [The Data Mesh as the end of the Datalake as we know it](https://www.wearedevelopers.com/videos/156-the-data-mesh-as-the-end-of-the-datalake-as-we-know-it) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)