From 8136c08d6a6a9a7d3c8a08d67ea7156e3c68e9fb Mon Sep 17 00:00:00 2001 From: Florian Berger Date: Thu, 17 Sep 2026 10:47:31 +0200 Subject: [PATCH 1/2] Add AI Discovery Radar to related resources --- README.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/README.md b/README.md index 320fa71..8da272b 100644 --- a/README.md +++ b/README.md @@ -58,6 +58,8 @@ file on-the-fly. - [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0. +- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): open dataset on whether the AI opt-out and discovery files exist in the first place, measured monthly on a rotating sample of ~10,700 domains drawn from a fixed frame, so the same routes can be compared over time. In the September 2026 run `robots.txt` was present on 84.7% of observed hosts, while the dedicated opt-out formats each stayed below 1%. Aggregates only, no per-domain results. CSV/JSON, CC BY 4.0, with a published ruleset and a DOI. + ## Contributing A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`. From 9f039df105a9424ce702d213efbb89fca4845eef Mon Sep 17 00:00:00 2001 From: Florian Berger Date: Thu, 17 Sep 2026 14:01:37 +0200 Subject: [PATCH 2/2] Shorten entry: drop dated figures, spell out the file types Updated the description of the AI Discovery Radar dataset to clarify its purpose and measurement methodology. --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 8da272b..4ff009d 100644 --- a/README.md +++ b/README.md @@ -58,7 +58,7 @@ file on-the-fly. - [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0. -- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): open dataset on whether the AI opt-out and discovery files exist in the first place, measured monthly on a rotating sample of ~10,700 domains drawn from a fixed frame, so the same routes can be compared over time. In the September 2026 run `robots.txt` was present on 84.7% of observed hosts, while the dedicated opt-out formats each stayed below 1%. Aggregates only, no per-domain results. CSV/JSON, CC BY 4.0, with a published ruleset and a DOI. +- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0. ## Contributing