Data Engineer - Web Scraping

Jobgether
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate

Job location

Tech stack

HTML
JavaScript
API
Airflow
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Big Data
Cloud Computing
Computer Programming
Databases
Continuous Integration
Data Cleansing
Information Engineering
Web Scraping
Data Integrity
Data Systems
Data Warehousing
Fiddler (Software)
Python
Cloud Services
Selenium
Software Engineering
SQL Databases
Workflow Management Systems
XPath
Data Processing
Data Ingestion
Postman
Pandas
Gitlab-ci
Kubernetes
Information Technology
Web Technologies
Amazon Web Services (AWS)
Data Pipelines
Docker
Jenkins

Job description

This role is an exciting opportunity to build scalable web scraping solutions and data pipelines that power business-critical insights. You will work closely with analysts, engineers, and cross-functional teams to develop reliable datasets that support strategic decision-making. With significant ownership over your projects, you'll design, automate, and optimize data collection workflows while ensuring data quality and reliability. The environment encourages innovation, collaboration, and continuous learning, giving you the freedom to experiment with new technologies and improve existing processes. If you enjoy solving complex data challenges and building automation at scale, this role offers an excellent platform to make a meaningful impact. Accountabilities:

As a Data Engineer specializing in web scraping, you will design, develop, and maintain automated data collection systems while ensuring the quality, accuracy, and availability of large-scale datasets. You will collaborate across teams to build efficient, reliable, and scalable data solutions.

  • Collaborate with analysts and stakeholders to understand data requirements and deliver tailored data solutions.
  • Design, develop, and maintain web scrapers for a wide range of structured and unstructured data sources.
  • Clean, transform, validate, and manipulate large datasets using Python and Pandas.
  • Build and maintain data ingestion pipelines into databases or data warehouses.
  • Schedule, monitor, and optimize scraping workflows using orchestration tools such as Apache Airflow.
  • Develop quality control checks to ensure data integrity, consistency, and availability.
  • Investigate and resolve data pipeline issues and time-sensitive production incidents.
  • Design and enhance internal tools, automation frameworks, and platform capabilities to improve operational efficiency.
  • Work closely with cross-functional engineering teams to implement scalable and maintainable data processing workflows.

Requirements

The ideal candidate combines strong software engineering skills with hands-on experience in web scraping, data processing, and automation. Success in this role requires both technical expertise and a proactive, problem-solving mindset.

  • Bachelor's or Master's degree in Computer Science or a related technical discipline.

  • 2-4 years of professional software development experience.

  • Strong programming skills in Python and solid SQL/database knowledge.

  • Advanced experience using the Pandas library for data cleaning, transformation, and analysis.

  • Experience working with web technologies, including HTML, JavaScript, APIs, and related protocols.

  • Proven experience processing, cleaning, and transforming large datasets.

  • Familiarity with web scraping frameworks and tools such as Selenium, Scrapy, XPath, Fiddler, or Postman.

  • Experience with workflow orchestration tools such as Apache Airflow or similar platforms.

  • Knowledge of Docker containerization; Kubernetes experience is an advantage.

  • Familiarity with CI/CD tools such as Jenkins or GitLab CI/CD.

  • Experience working with cloud services, particularly AWS technologies such as S3, RDS, Lambda, SNS, or SQS, is preferred.

  • Strong analytical thinking, attention to detail, communication skills, and a passion for automation and continuous improvement.

  • Opportunity to work on challenging projects supporting a leading global asset management environment.

  • High level of ownership and autonomy in a collaborative, team-oriented culture.

  • Exposure to modern data engineering, web scraping, cloud, and automation technologies.

  • Collaborative environment with experienced engineering, product, and data professionals.

  • Opportunities for continuous learning, professional growth, and technical skill development.

  • Merit-driven culture that values innovation, initiative, and individual contributions.

  • Flexible, technology-focused environment with opportunities to work on impactful data products.

Apply for this position