> Markdown version of [/jobs/ext/10115-data-engineer-link-key-people](https://www.wearedevelopers.com/jobs/ext/10115-data-engineer-link-key-people). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - Link Key People - **Company:** Veeva Systems - **Location:** Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Architectural Patterns, Cloud Engineering, Data Architecture, Information Engineering, Data Integrity, Data Structures, Python (Programming Language), Machine Learning, Data Logging, Large Language Models, Apache Spark, Data Lakes, Pyspark, Solid Principles, Veeva, Data Pipelines - **Published:** May 20, 2026 - **Apply:** https://es.indeed.com/viewjob?jk=a1750d9fabf4a23e ## About the Role Do you have experience in Spark?, Do you have a Master's degree?, * Experience with Python and Apache Spark/PySpark to handle massive datasets * Expertise in building cloud-native software within AWS or GCP * Background in designing and maintaining modern architectures, specifically Data Lakes, Lakehouses and Warehouses (DeltaLake, Redshift) * Experience operating LLM systems in production, including third-party model providers, human/data feedback loops, and multi-model traffic orchestration * Driving technical execution within Agile environments, utilizing strong English communication skills to align with global stakeholders ## Description At Veeva Link, we're building the intelligence layer for life sciences, creating connected data applications that accelerate drug development and significantly improve patient outcomes. Our core belief: combining the highest quality data with state-of-the-art software delivers immense value. We believe execution matters most. Progress comes from speed, accuracy, and quality in what we build every day. Our engineering approach emphasizes clear product definitions leading to meticulous technical designs. We leverage a modern tech stack to deliver inherently reliable applications and a great user experience. As a Data Engineer, you will be responsible for the life cycle of the data that defines the healthcare landscape. You will design, build, and maintain the robust data pipelines required to ingest and process global Healthcare Organization (HCO) data. You will be a key architect in managing the complex hierarchical relationships of over 4 million entities, ensuring data integrity, scalability, and seamless delivery to downstream stakeholders. We are an AI-forward team. We actively promote and integrate AI-driven development practices, leveraging LLMs and automation to accelerate our coding, optimize our pipelines, and stay at the forefront of the rapidly evolving data landscape. We expect you to embrace these technologies to enhance both your productivity and the product's capabilities. What You'll Do * You will architect the "Data DNA" used by global biopharmas to power AI driven insights and accelerate medical innovation * You will be at the forefront of enabling global AI initiatives within high-stakes, high-impact environments * You will design PySpark pipelines and collaborate on ML models to integrate diverse data into robust Lakehouse architecture * Identify, implement, and maintain end-to-end HCO data pipelines * Refine data structures and processing logic to meet the rapidly changing demands of the market * Deploy solid principles and clean patterns to data engineering tasks * Advance the long-term architectural roadmap * Validate high quality and availability of HCO deliveries * Govern and optimize the underlying infrastructure * Own monitoring, logging, and performance metrics from day one * A proactive interest in using AI tools to streamline development and solve complex data problems ## Related Videos - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [The Data Mesh as the end of the Datalake as we know it](https://www.wearedevelopers.com/videos/156-the-data-mesh-as-the-end-of-the-datalake-as-we-know-it) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)