Senior Software Engineer (Web Scraping, Data Acquisition, Data Collection) - Nielseniq

Jobrapido
Barcelona, Spain
5 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Application Programming Interfaces (APIs) Databases Data Centers Web Scraping Software Debugging Python (Programming Language) MongoDB Redis Reverse Engineering Selenium Session Management
+6 more
Parquet Influxdb Playwright Avro Mitmproxy Stream Processing

Job description

OverviewHaga clic en “Solicitar” a continuación para enviar su candidatura.Asegúrese de que su CV está actualizado y de que ha leído primero las especificaciones del puesto.In this role you will design and operate scalable distributed scraping systems to extract high-quality data from thousands of domains.You will work within the Scraping and R&D team to push innovative approaches and share knowledge, collaborating with Customer Success, Operations, and Data.You’ll tackle anti-bot challenges and build monitoring to prevent blockers, contributing to fast, scalable visibility for retailers and brands.This is a hands-on, problem?solving role for someone who loves turning difficult data problems into reliable solutions.Compensaciones / BeneficiosFlexible working environmentVolunteer time offLinkedIn LearningEmployee-Assistance-Program (EAP)ResponsabilidadesDesign and build scalable distributed scraping architectures for thousands of domainsReverse engineer APIs and mobile apps using tools like Frida, mitmproxy, Charles Proxy, burpDetect and bypass anti-bot measures to create human-like crawlersLead technical innovation in scraping, prototype approaches, share knowledge with the teamSet up anomaly detection and monitoring to catch blockers or failures earlyRequisitos principales5+ years of Python experience and involvement in large-scale scraping projectsDeep understanding of anti-bot technologies and bypass techniques, fingerprinting, WAFs, JavaScript challengesExperience with Scrapy, Requests, httpx, Selenium, Playwright or similar headless driversProficiency with proxies (residential, datacenter, rotating), session management, cookie injection, TLS tweakingStrong debugging skills across browser, HTTP traffic, xqziphu and device-level interactionsIndependent problem solver with a track record of solving new problemsFamiliarity with Parquet, Avro or real-time stream processingKnowledge of databases like MongoDB, InfluxDB, and Redis and how to optimize themIndependent thinkingStrong problem-solving abilitiesSelf-motivation and proactivityPython for large-scale scrapingAnti-bot bypass techniques, fingerprinting, WAFs, JavaScript challengesScrapy, Requests, httpx, Selenium, Playwright or headless drivers

Requirements

Requisitos principales5+ years of Python experience and involvement in large-scale scraping projects Deep understanding of anti-bot technologies and bypass techniques, fingerprinting, WAFs, JavaScript challenges Experience with Scrapy, Requests, httpx, Selenium, Playwright or similar headless drivers Proficiency with proxies (residential, datacenter, rotating), session management, cookie injection, TLS tweaking Strong debugging skills across browser, HTTP traffic, xqziphu and device-level interactions Independent problem solver with a track record of solving new problems Familiarity with Parquet, Avro or real-time stream processing Knowledge of databases like MongoDB, InfluxDB, and Redis and how to optimize them Independent thinking Strong problem-solving abilities Self-motivation and proactivity Python for large-scale scraping Anti-bot bypass techniques, fingerprinting, WAFs, JavaScript challenges Scrapy, Requests, httpx, Selenium, Playwright or headless drivers

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:54 min

The technical evolution of modern web scraping infrastructure

Chris Heilmann Chris Heilmann +4 · LIVE

2:01 min

Migrating existing applications from MongoDB to Postgres

Nikita Shamgunov Nikita Shamgunov · World Congress 2024

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

1:21 min

Realizing the limitations of MongoDB for live statistics

Josip Stuhli Josip Stuhli · World Congress 2023

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all