mirror of
https://github.com/ai-robots-txt/ai.robots.txt.git
synced 2026-09-28 04:44:29 +02:00
Shorten entry: drop dated figures, spell out the file types
Updated the description of the AI Discovery Radar dataset to clarify its purpose and measurement methodology.
This commit is contained in:
parent
8136c08d6a
commit
9f039df105
1 changed files with 1 additions and 1 deletions
|
|
@ -58,7 +58,7 @@ file on-the-fly.
|
|||
|
||||
- [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0.
|
||||
|
||||
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): open dataset on whether the AI opt-out and discovery files exist in the first place, measured monthly on a rotating sample of ~10,700 domains drawn from a fixed frame, so the same routes can be compared over time. In the September 2026 run `robots.txt` was present on 84.7% of observed hosts, while the dedicated opt-out formats each stayed below 1%. Aggregates only, no per-domain results. CSV/JSON, CC BY 4.0, with a published ruleset and a DOI.
|
||||
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0.
|
||||
|
||||
## Contributing
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue