Data Engineer - Web Scraping

Jobgether
Álava, Spain
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

HTML JavaScript (Programming Language) Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon S3 Apache HTTP Server Big Data Cloud Computing Computer Programming Databases Continuous Integration
+27 more
Data Cleansing Information Engineering Web Scraping Data Integrity Data Systems Data Warehousing Fiddler (Software) Python (Programming Language) Cloud Services Selenium Software Engineering SQL Databases Unstructured Data Workflow Management Systems XPath Data Processing Data Ingestion Postman Pandas Gitlab-ci Kubernetes Information Technology Web Technologies Amazon Simple Queue Service (SQS) Data Pipelines Docker Jenkins

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Engineer - Web Scraping based in Spain. This role is an exciting opportunity to build scalable web scraping solutions and data pipelines that power business-critical insights. You will work closely with analysts, engineers, and cross-functional teams to develop reliable datasets that support strategic decision-making. With significant ownership over your projects, you’ll design, automate, and optimize data collection workflows while ensuring data quality and reliability. The environment encourages innovation, collaboration, and continuous learning, giving you the freedom to experiment with new technologies and improve existing processes. If you enjoy solving complex data challenges and building automation at scale, this role offers an excellent platform to make a meaningful impact.AccountabilitiesAs a Data Engineer specializing in web scraping, you will design, develop, and maintain automated data collection systems while ensuring the quality, accuracy, and availability of large-scale datasets. You will collaborate across teams to build efficient, reliable, and scalable data solutions.Collaborate with analysts and stakeholders to understand data requirements and deliver tailored data solutions.Design, develop, and maintain web scrapers for a wide range of structured and unstructured data sources.Clean, transform, validate, and manipulate large datasets using Python and Pandas.Build and maintain data ingestion pipelines into databases or data warehouses.Schedule, monitor, and optimize scraping workflows using orchestration tools such as Apache Airflow.Develop quality control checks to ensure data integrity, consistency, and availability.Investigate and resolve data pipeline issues and time-sensitive production incidents.Design and enhance internal tools, automation frameworks, and platform capabilities to improve operational efficiency.Work closely with cross-functional engineering teams to implement scalable and maintainable data processing workflows.RequirementsBachelor’s or Master’s degree in Computer Science or a related technical discipline.2-4 years of professional software development experience.Strong programming skills in Python and solid SQL/database knowledge.Advanced experience using the Pandas library for data cleaning, transformation, and analysis.Experience working with web technologies, including HTML, JavaScript, APIs, and related protocols.Proven experience processing, cleaning, and transforming large datasets.Familiarity with web scraping frameworks and tools such as Selenium, Scrapy, XPath, Fiddler, or Postman.Experience with workflow orchestration tools such as Apache Airflow or similar platforms.Knowledge of Docker containerization; Kubernetes experience is an advantage.Familiarity with CI/CD tools such as Jenkins or GitLab CI/CD.Experience working with cloud services, particularly AWS technologies such as S3, RDS, Lambda, SNS, or SQS, is preferred.Strong analytical thinking, attention to detail, communication skills, and a passion for automation and continuous improvement.BenefitsOpportunity to work on challenging projects supporting a leading global asset management environment.High level of ownership and autonomy in a collaborative, team-oriented culture.Exposure to modern data engineering, web scraping, cloud, and automation technologies.Collaborative environment with experienced engineering, product, and data professionals.Opportunities for continuous learning, professional growth, and technical skill development.Merit-driven culture that values innovation, initiative, and individual contributions.Flexible, technology-focused environment with opportunities to work on impactful data products.#J-*****-Ljbffr

Requirements

Bachelor’s or Master’s degree in Computer Science or a related technical discipline. 2-4 years of professional software development experience. Strong programming skills in Python and solid SQL/database knowledge. Advanced experience using the Pandas library for data cleaning, transformation, and analysis. Experience working with web technologies, including HTML, JavaScript, APIs, and related protocols. Proven experience processing, cleaning, and transforming large datasets. Familiarity with web scraping frameworks and tools such as Selenium, Scrapy, XPath, Fiddler, or Postman. Experience with workflow orchestration tools such as Apache Airflow or similar platforms. Knowledge of Docker containerization; Kubernetes experience is an advantage. Familiarity with CI/CD tools such as Jenkins or GitLab CI/CD. Experience working with cloud services, particularly AWS technologies such as S3, RDS, Lambda, SNS, or SQS, is preferred. Strong analytical thinking, attention to detail, communication skills, and a passion for automation and continuous improvement.

Benefits & conditions

Opportunity to work on challenging projects supporting a leading global asset management environment. High level of ownership and autonomy in a collaborative, team-oriented culture. Exposure to modern data engineering, web scraping, cloud, and automation technologies. Collaborative environment with experienced engineering, product, and data professionals. Opportunities for continuous learning, professional growth, and technical skill development. Merit-driven culture that values innovation, initiative, and individual contributions. Flexible, technology-focused environment with opportunities to work on impactful data products. #J-*****-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:18 min

Identifying deceptive server content and parsing messy HTML

Vidas Bacevičius Vidas Bacevičius · WWC 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:21 min

Projecting external HTML content using default and named slots

Rowdy Rabouw Rowdy Rabouw · WWC 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

3:19 min

Automating data extraction pipelines using natural language assistants

Vidas Bacevičius Vidas Bacevičius · WWC 2025

Videos

See all

Related articles

See all