Research Engineer - Web Crawlers

Eleven Labs Inc.
Poland, ME, United States
7 days ago
Apply on www.workingnomads.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Training Data HTML Artificial Intelligence Data Deduplication Distributed Systems Github Machine Learning Web Crawlers Kubernetes

Job description

  • Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
  • Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
  • Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
  • Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
  • Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.
  • Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend., We are looking for a Research Engineer to join the research team at ElevenLabs, focused on large-scale web crawling for our frontier AI models. The quality of our models is bounded by the quality and scale of the data behind them, and you will own the crawling systems that source world-class data from the open web. You will thrive in this role if you enjoy:
  • Building and operating large-scale, distributed web crawlers that discover, fetch, and extract data across billions of pages reliably and efficiently.
  • Solving hard crawling problems such as content extraction from messy HTML, deduplication at web scale, freshness and recrawl strategies, and politeness and rate-limit handling.
  • Designing targeted crawling pipelines that find high-value data sources, including audio, video, and multilingual content, and turn them into clean training-ready datasets.
  • Creating tooling and infrastructure that lets researchers request, monitor, and explore newly crawled web data quickly and reliably.

Requirements

We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:

  • Hands-on experience building and scaling web crawlers or scraping systems, ideally in support of machine learning training data.
  • Strong engineering skills in distributed systems at scale (e.g., Kubernetes, queue-based architectures, or custom pipelines processing billions of documents).
  • The capacity to autonomously evaluate the quality, coverage, and compliance of crawled data, and to build the tooling to measure it.

About the company

ElevenLabs is an AI research and product company transforming how we interact with technology.

We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world’s most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We’ve raised $781M in funding and our last valuation was $11B - multiples of 11, always. We have expanded from voice into three main platforms:

  • ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.
  • ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.
  • ElevenAPI gives developers access to our leading AI audio foundational models.

Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you., This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.workingnomads.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:20 min

Securing competitive intelligence using automated web extraction crawlers

Coffee With Developers

2:21 min

Projecting external HTML content using default and named slots

Rowdy Rabouw Rowdy Rabouw ¡ World Congress 2022

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters ¡ World Congress 2023

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter ¡ World Congress 2022

1:19 min

Tailoring custom web scrapers for artificial intelligence training

Vidas Bacevičius Vidas Bacevičius · World Congress 2025

6:12 min

Streaming HTML content natively using declarative processing instructions

Chris Heilmann +2 ¡ LIVE

Videos

See all

Related articles

See all