Web Crawler engineer

ONE STOP COLLECTIBLE CORP
San Francisco, CA, United States
5 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

JavaScript (Programming Language) Dynamic Content

Job description

  • Build a distributed crawler that can handle 100M+ pages per day
  • Optimize crawl politeness and rate limiting across thousands of domains
  • Design systems to detect and handle dynamic content, JavaScript rendering, and anti-bot measures
  • Create intelligent crawl scheduling and prioritization algorithms for maximum coverage efficiency

Requirements

  • You have extensive experience building and scaling web crawlers, or would be excited to ramp up very quickly
  • You have experience with some high performance language (C++, Rust, etc.)
  • You are familiar with TypeScript, Playwright, modern web design, CDP (Chrome DevTools Protocol)
  • You’re comfortable optimizing a system to an exceptional degree
  • You care about the problem of finding high quality knowledge and recognize how important this is for the world

Benefits & conditions

  • Location: This is an in-person opportunity in San Francisco.
  • Visas: We’re happy to sponsor international candidates (e.g., STEM OPT, OPT, H1B, O1, E3). While we cannot guarantee your visa, we have historically been successful in sponsoring candidates from all over the world. If you receive an offer, our team will work hard to get you a visa.
  • Benefits: We offer premium healthcare benefits (medical, dental, vision), fertility benefits, 16 weeks of fully paid parental leave for all new parents, and a monthly wellness stipend to all of our employees.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:20 min

Securing competitive intelligence using automated web extraction crawlers

Coffee With Developers

3:09 min

The history of JavaScript and standardizing browser compatibility

Sasha Shynkevich · LIVE

1:46 min

Building reactive content with dynamic script tags

André Dietrich André Dietrich · World Congress 2024

1:57 min

Demonstrating web data collection using JavaScript browser automation libraries

Tim Ruscica · Coffee With Developers

2:37 min

Addressing the complaints about JavaScript as a programming language

Sasha Shynkevich · LIVE

1:50 min

Causes of cumulative layout shift and screen instability

Nicolas Frizzarin Nicolas Frizzarin · World Congress 2024

Videos

See all

Related articles

See all