ScrubberDuck working with public machine-readable feeds

Public Crawler Feeds

A consistent public observation surface for search and AI bots

Why these feeds exist

Scrubnet publishes authorised source-derived HTML, JSON, TXT and Markdown feeds so crawler behaviour can be observed on a consistent public surface. Stable structures, source links and timestamps make it easier to describe discovery, recrawling and request distribution across search, AI and archive bots.

Questions we observe

Explore the data

Start with /feed/sitemap.xml to discover active feeds and their <lastmod> values. Then use the live logs dashboard to filter requests by path, bot, verification status and date, or compare requests by crawler and format.

Technical access

Methodological limits

Access logs confirm that a request reached a Scrubnet feed. They do not prove that content was indexed, ranked, used for model training, retrieved in an answer or cited to a user. Findings are descriptive observations from Scrubnet’s environment and do not establish why a crawler behaved in a particular way.

ScrubberDuck

ScrubberDuck/1.0 is the robots-aware, low-rate feed compiler that creates and refreshes authorised participating-site feeds. It supplies the observation surface; it is not a search engine or ranking crawler.

ScrubberDuck, Scrubnet’s research crawler

Research and crawler teams

To contribute a site, suggest an observational question or discuss access to supporting data, email contact@scrubnet.org.