Shorten entry: drop dated figures, spell out the file types

Updated the description of the AI Discovery Radar dataset to clarify its purpose and measurement methodology.
This commit is contained in:
Florian Berger 2026-09-17 14:01:37 +02:00 • committed by GitHub
commit 9f039df105
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -58,7 +58,7 @@ file on-the-fly.
- [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0.
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): open dataset on whether the AI opt-out and discovery files exist in the first place, measured monthly on a rotating sample of ~10,700 domains drawn from a fixed frame, so the same routes can be compared over time. In the September 2026 run `robots.txt` was present on 84.7% of observed hosts, while the dedicated opt-out formats each stayed below 1%. Aggregates only, no per-domain results. CSV/JSON, CC BY 4.0, with a published ruleset and a DOI.
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0.
## Contributing