What is Scrubnet?
Scrubnet is an independent public observatory for understanding how search crawlers and AI bots behave. We compile authorised website content into machine-readable feeds, record how verified bots discover and revisit them, and publish the resulting data, observations and practical tools.
Why study crawlers?
- 🤖 Search engines, AI assistants and answer engines use different crawlers with different purposes.
- 📊 Their real fetch behaviour is often hidden behind assumptions, documentation and incomplete third-party data.
- 🔭 Scrubnet provides a consistent observation surface for discovery, recrawling, freshness and request distribution by format.
Who It’s For
- 🔎 SEO and GEO practitioners: use real crawler data to inform audits, analysis and recommendations.
- 🧪 Researchers and crawler teams: examine public feeds, fetch patterns and technical findings.
- 🌐 Site owners: contribute a site for free and help broaden the research dataset.
Our Principles
- 🛡️ Independent: our work is not tied to a search engine, AI platform or vendor.
- 🔬 Evidence-led: we separate observed behaviour from hypotheses and marketing claims.
- 🌍 Open by default: feeds, live logs and findings are public wherever access and privacy allow.
What Scrubnet brings together
Scrubnet combines a growing set of machine-readable feeds, a public crawler log dashboard, observational reports and free tools. Together they help us move beyond crawler speculation and build a clearer picture of how verified crawlers interact with public content.
Meet ScrubberDuck
ScrubberDuck is our lightweight, robots-aware feed compiler. It collects authorised public content and creates the Scrubnet feeds used to observe discovery, formats, freshness signals and recrawl behaviour.
It’s designed to minimise load, avoid unnecessary requests, and respect robots.txt.
If you see ScrubberDuck in your logs, it means your site is contributing authorised content to the Scrubnet observatory.
User-Agent: ScrubberDuck/1.0 (+https://scrubnet.org)
How the research works
We add authorised participating websites, fetch their public pages efficiently and publish consistent machine-readable feeds. We then monitor which verified and unverified bots request those feeds, when they return, which formats they choose and how requests relate to content changes.
The observations feed into descriptive reports, technical SEO and GEO guidance, and tools such as SEO Scrubbox. Adding more varied sites broadens the observable content base and makes the dataset more representative of the participating sites.
Participation is free: add a website with up to 50,000 public URLs. There are no visibility or ranking guarantees.
Crawlers we monitor
The live dashboard records recognised search, AI and archive crawlers requesting Scrubnet feeds, including:
- Googlebot – Google Search and Discover
- GPTBot – OpenAI’s web crawler
- ClaudeBot – Anthropic’s crawler for Claude
- PerplexityBot – Perplexity AI’s research assistant bot
- bingbot – Microsoft Bing search crawler
- BingPreview – Bing page preview bot
- CCBot – Common Crawl archive bot
- DuckDuckBot – DuckDuckGo crawler
- Applebot – Apple Siri and Spotlight crawler
Get Involved
Add a site to expand the observational dataset, explore the live logs, or collaborate with us on a crawler analysis or case study.
Or reach out at contact@scrubnet.org
SEO Scrubbox
A Chrome extension for technical SEO and crawler diagnostics. Compare view-source vs rendered signals, spot canonical drift, validate JSON-LD, audit sitemaps, hreflang, redirects, and crawl signals without leaving the page.