> Markdown version of [/videos/502-shoot-for-the-moon-machine-learning-for-automated-online-ad-detection?t=1124](https://www.wearedevelopers.com/videos/502-shoot-for-the-moon-machine-learning-for-automated-online-ad-detection?t=1124). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Shoot for the moon - machine learning for automated online ad detection When computer vision proved too heavy for real-time ad blocking, engineers pivoted to DOM analysis. Discover how they achieved sub-millisecond inference in-browser using XGBoost and TensorFlow.js. - **Speakers:** Humera Minhas, Parinitha Hirehal - **Event:** World Congress 2022 - **Published:** June 15, 2022 - **Duration:** 31:23 - **URL:** https://www.wearedevelopers.com/videos/502-shoot-for-the-moon-machine-learning-for-automated-online-ad-detection ## Summary Traditional ad filtering relies on static, manually maintained filter lists like EasyList, which demand massive human intervention and struggle against evolving ad circumvention. To address this bottleneck, "Project Moonshot" replaces rigid rules with machine learning to automate online ad detection. While initial experiments relied on perceptual ad detection using heavy computer vision models like YOLO and ResNet, these proved too resource-intensive for real-time browser environments. The team successfully pivoted to analyzing HTML page structures, extracting critical data from meta tags, URLs, and DOM trees to identify ad components efficiently. To build a robust training pipeline, the engineers developed a custom crawler using Headless Chrome and Puppeteer, leveraging a modified version of Adblock Plus to label rather than block content. This approach generated a massive ground-truth dataset of two million nodes. The team immediately faced severe data imbalances—where ads constituted only five percent of the dataset—which they resolved through strategic data augmentation. Additionally, optimizing their preprocessing architecture reduced the time needed to convert HTML into adjacency and feature matrices from three days down to just three minutes. During model benchmarking, they found that while standard Graph Neural Networks (GNNs) struggled with subgraph processing inefficiencies and yielded F1 scores around 30%, tree-based classifiers like XGBoost applied to node features achieved impressive F1 scores between 75% and 90%. Deploying ML models directly into the browser required converting complex Python research into a lightweight JavaScript implementation. A major architectural breakthrough involved moving the model inference from the frequently reloaded content script to the persistent background script. This optimization drastically reduced latency, achieving an inference speed of 0.3 milliseconds per node using TensorFlow.js. By building a scalable, in-browser classification system, the team established a robust defense against ad publishers—where any successful circumvention effort ultimately just feeds new training data back into the system to further harden the model. **Keywords:** machine learning ad detection, html structure analysis, automated filter lists, graph neural networks, browser model deployment, puppeteer web crawling, xgboost node classification, tensorflow.js inference, imbalanced data augmentation, dom tree feature extraction, background script optimization, self-supervised representation learning, ad blocking circumvention, node embeddings ## Chapters 1. **Using machine learning to automate online advertisement detection** (00:05) — Project Moonshot aims to filter intrusive internet ads using artificial intelligence and machine learning. 1. **Understanding the limitations of manual ad filter lists** (02:06) — Traditional filter lists demand significant human intervention to manually review and update rules against circumvention. 1. **Establishing an experimentation pipeline for ad detection models** (04:00) — The end-to-end process involves data collection, pre-processing, model benchmarking, and final deployment. 1. **Choosing HTML structures over computer vision for ad data** (05:07) — Analyzing HTML tags and page structure provides better performance than using heavy computer vision models. 1. **Generating ground truth data at scale using web crawlers** (07:16) — A custom crawler combined with headless Chrome and Adblock Plus creates large-scale labeled datasets. 1. **Pre-processing raw HTML into feature and adjacency matrices** (09:10) — Converting HTML into a JSON tree structure enables the generation of matrices for machine learning input. 1. **Overcoming challenges with unbalanced data and pipeline speed** (11:44) — Data augmentation resolves highly unbalanced distributions while server upgrades drastically reduce processing times. 1. **Applying graph neural networks for HTML node classification** (14:03) — Message passing mechanisms generate node embeddings that allow classifiers to distinguish ads from legitimate content. 1. **Evaluating model performance and utilizing self-supervised learning** (18:44) — Tree-based models like XGBoost achieve better F1 scores than direct graph neural networks for complex web graphs. 1. **Deploying Python machine learning models into JavaScript environments** (23:20) — Migrating models to background scripts using TensorFlow.js reduces inference latency and prevents continuous reloading. 1. **Addressing model circumvention and future machine learning objectives** (26:22) — Continuous data collection from ad publishers attempting to bypass detection further trains and hardens the model. ## Related Moments - [Real-world implementations of browser-based machine learning](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) (from "Machine learning in the browser with TensorFlowjs") - [Integrating embedded machine learning models into the Firefox browser](https://www.wearedevelopers.com/videos/1779-wearedevelopers-live-ai-and-privacy-on-device-ai-and-more) (from "WeAreDevelopers LIVE – AI and Privacy, On-Device AI and More") - [Implementing predictive prefetch with machine learning](https://www.wearedevelopers.com/videos/287-state-of-angular) (from "State of Angular") - [Leveraging Chrome AI and nano models for web applications](https://www.wearedevelopers.com/videos/1770-generate-ai-in-the-browser-with-chrome-ai-raymond-camden) (from "Generate AI in the Browser with Chrome AI - Raymond Camden") - [Evaluating page performance using customized ML prediction models](https://www.wearedevelopers.com/videos/1771-ai-is-an-electric-bike-for-the-brain-stoyan-stefanov) (from "AI is an Electric Bike for the Brain - Stoyan Stefanov") - [Challenges with open source monetization and AI scraping](https://www.wearedevelopers.com/videos/1786-wearedevelopers-live-ai-freelancing-keeping-up-with-tech-and-more) (from "WeAreDevelopers LIVE – AI, Freelancing, Keeping Up with Tech and More") ## Related Articles - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [The Web We Broke (And Why AI Agents Are Paying the Price) - AgentCon Berlin](https://www.wearedevelopers.com/magazine/735-the-web-we-broke-and-why-ai-agents-are-paying-the-price-agentcon-berlin) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) ## Related Jobs - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/2725456-data-scientist) at **Almedia** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [ML Engineer](https://www.wearedevelopers.com/jobs/48448-ml-engineer) at **Docker, Inc.** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Staff ML Engineer](https://www.wearedevelopers.com/jobs/48463-staff-ml-engineer) at **Docker, Inc.** - [Model Implementation Engineer](https://www.wearedevelopers.com/jobs/48421-model-implementation-engineer) at **Sciforium**