Data Engineer

Veeva Link
Barcelona, Spain
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
€65,000.0 - €110,000.0
Working hours
Regular working hours

Tech stack

Agile Methodology Artificial Intelligence Amazon Web Services Architectural Patterns Cloud Engineering Information Engineering Data Integrity Data Structures Python (Programming Language) Machine Learning Data Logging Large Language Models
+6 more
Apache Spark Data Lakes Pyspark Solid Principles Veeva Data Pipelines

Job description

At Veeva Link, we are building the intelligence layer for life sciences, creating connected data applications that accelerate drug development and significantly improve patient outcomes. Our core belief is that combining the highest quality data with state-of-the-art software delivers immense value.

As a Data Engineer, you will be responsible for the life cycle of the data that defines the healthcare landscape. You will design, build, and maintain the robust data pipelines required to ingest and process global Healthcare Organization (HCO) data. You will be a key architect in managing the complex hierarchical relationships of over 4 million entities, ensuring data integrity, scalability, and seamless delivery to downstream stakeholders.

We are an AI-forward team and actively promote AI-driven development practices, leveraging LLMs and automation to accelerate coding, optimize pipelines, and stay at the forefront of the evolving data landscape.

  • You will architect the “Data DNA” used by global biopharmas to make data-driven decisions
  • You will be at the forefront of enabling global AI initiatives within high-stakes, high-impact environments
  • You will design PySpark pipelines and collaborate on ML models to integrate diverse data into a robust lakehouse architecture
  • Identify, implement, and maintain end-to-end HCO data pipelines
  • Refine data structures and processing logic to meet the rapidly changing demands of the market
  • Deploy solid principles and clean patterns to data engineering tasks
  • Advance the long-term architectural roadmap
  • Validate high quality and availability of HCO deliveries
  • Govern and optimize the underlying infrastructure
  • Own monitoring, logging, and performance metrics from day one
  • A proactive interest in using AI tools to streamline development and solve complex data problems

Requirements

  • Experience with Python and Apache Spark/PySpark to handle massive datasets
  • Expertise in building cloud-native software within AWS or GCP
  • Background in designing and maintaining modern architectures, specifically Data Lakes, lakehouses and warehouses (DeltaLake, Redshift)
  • Experience operating LLM systems in production, including third-party model providers, human/data feedback loops, and multi-model traffic orchestration
  • Driving technical execution within Agile environments, utilizing strong English communication skills to align with global stakeholders

Benefits & conditions

  • Comprehensive benefits package
  • Fitness reimbursement
  • Veeva Work Anywhere

Total compensation: 65,000 - 110,000 EUR.

The total compensation range listed here has been provided to comply with local regulations and represents a potential total compensation range for this role. Please note that actual total compensation may vary within the range above or below, depending on experience and location. We look at compensation for each individual and base our offer on your unique qualifications, experience, and expected contributions. This position’s total compensation may be composed of other types of compensation, such as variable bonus and/or stock bonus.

As an equal opportunity employer, Veeva is committed to fostering a culture of inclusion and growing a diverse workforce. Diversity makes us stronger. It comes in many forms. Gender, race, ethnicity, religion, politics, sexual orientation, age, disability and life experience shape us all into unique individuals. We value people for the individuals they are and the contributions they can bring to our teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobleads.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:24 min

The governance failures of centralized data lakes

Mario Meir-Huber · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

6:24 min

Distributed data lakes and containerized computing clusters

Ulrich Wurstbauer +1 · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all