Commit graph

736 commits

Author SHA1 Message Date
ai.robots.txt
2acefa38cc Merge pull request #274 from andyman01/add-qwen-ernie-doubao-mistralindex v1.52
Add four AI-lab crawlers named by Dutch news publishers
2026-09-07 03:41:26 +00:00
Glyn Normington
a595115206
Merge pull request #274 from andyman01/add-qwen-ernie-doubao-mistralindex
Add four AI-lab crawlers named by Dutch news publishers
2026-09-07 04:41:12 +01:00
andyman01
48abe91f99 Add four AI-lab crawlers named by Dutch news publishers
QwenBot, ERNIEBot, DoubaoBot and MistralAI-Index appear in the robots.txt
of Dutch news sites, measured against 83 domains in August 2026, but were
missing from this list.

MistralAI-Index is documented by the vendor at https://docs.mistral.ai/robots
and is the third of Mistral's three crawlers.

For QwenBot, ERNIEBot and DoubaoBot the evidence is different in kind and
the fields say so: the operator is identifiable from the name and the
company's own model family, but none of the three is documented on a
vendor crawler page I could find, so "respect" stays "Unclear at this
time." They are included because publishers are already naming them,
which is the same reason this list carries other undocumented agents.

Only robots.json is changed; the other files are regenerated
automatically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
2026-09-06 09:37:40 +02:00
ai.robots.txt
f76020b4e4 Merge pull request #277 from andyman01/add-kimi-mistral-training
Add KimiBot, Kimi-SearchBot and MistralAI-Training
2026-09-06 03:57:23 +00:00
Glyn Normington
b509c62249
Merge pull request #277 from andyman01/add-kimi-mistral-training
Add KimiBot, Kimi-SearchBot and MistralAI-Training
2026-09-06 04:57:15 +01:00
ai.robots.txt
da24dd28d1 Merge pull request #281 from andyman01/add-oai-adsbot
Add OAI-AdsBot
2026-09-06 03:56:13 +00:00
Glyn Normington
397d77cd74
Merge pull request #281 from andyman01/add-oai-adsbot
Add OAI-AdsBot
2026-09-06 04:56:05 +01:00
andyman01
d37bb86786 Add KimiBot, Kimi-SearchBot and MistralAI-Training
All three are documented by the vendor itself:

- KimiBot and Kimi-SearchBot: https://www.kimi.ai/policies/kimi-crawlers
  The list already has Kimi-User; Moonshot documents three crawlers.
- MistralAI-Training: https://docs.mistral.ai/robots/
  The list already has MistralAI-User. Mistral documents three; the
  third, MistralAI-Index, is proposed separately in #274.

Only robots.json is changed, per review: the other files are regenerated
automatically. Rebased on main so that Diffbot-User, merged in the
meantime, is preserved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
2026-09-05 18:58:24 +02:00
andyman01
2d7931778c Add OAI-AdsBot
Added at a maintainer's request in #277, where I raised it but did not
propose it.

OpenAI documents it at https://developers.openai.com/api/docs/bots as the
fourth of its four crawlers. Two things in that entry are worth recording
in the fields rather than leaving to the reader:

- The page does not say OAI-AdsBot obeys robots.txt, and gives no
  disallow example, so "respect" is "Unclear at this time." rather than
  the "Yes" the other three OpenAI crawlers carry.
- OpenAI states the data is not used to train foundation models, and that
  the crawler only visits pages submitted as ads. What makes it belong on
  this list is the other half of the same paragraph: content from the
  landing page is also used to "determine when it's most relevant to show
  the ad to users".

Only robots.json is changed; the other files are regenerated
automatically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
2026-09-05 15:05:22 +02:00
Glyn Normington
f44d4386fc
Merge pull request #279 from saar-twito/main
Add AI Access Checker to additional resources
2026-09-04 17:00:28 +01:00
Glyn Normington
a28a34b049
Merge pull request #280 from ai-robots-txt/instructions
Disambiguate contribution instructions
2026-09-04 16:58:06 +01:00
Glyn Normington
984382f00f Disambiguate contribution instructions
Also fix some markdown issues.
2026-09-04 16:56:43 +01:00
Glyn Normington
282ef53f61
Merge pull request #276 from andyman01/add-diffbot-user
Add Diffbot-User
2026-09-04 16:45:26 +01:00
Saar Twito
538fc8de8f
Add AI Access Checker to additional resources 2026-09-04 12:11:49 +03:00
andyman01
ed6af83530 Add Diffbot-User
Diffbot documents two separate user agents. The list currently has only
the proactive crawler; this adds the user-triggered one.

From Diffbot's own documentation:

  Diffbot-User — This is used by requests originating on behalf of a
  human user browsing a URL using Diffbot software, in response to
  their input.

Source: https://docs.diffbot.com/docs/does-crawl-respect-robotstxt

Same distinction as GPTBot / ChatGPT-User and PerplexityBot /
Perplexity-User, so it is classified as an AI Assistant fetching only
when prompted by a user.

Observed in the robots.txt of five independent publishers in four
countries: nytimes.com, bbc.co.uk, lefigaro.fr, corriere.it and
usatoday.com.

Generated files updated with code/robots.py --convert.
2026-08-29 20:34:05 +02:00
Glyn Normington
7cd4c92343
Merge pull request #272 from peppe1337/de-access-index v1.51
Add KI-Zugangsindex to Related resources
2026-08-26 23:51:59 +01:00
Ideenschmiede CEO
5e73986001 Apply maintainer's suggestion: trim entry to avoid stale figures 2026-08-26 16:02:12 +00:00
Ideenschmiede CEO
3b67c9c54e Add KI-Zugangsindex to Related resources 2026-08-25 12:15:02 +00:00
Known Agents
0037a9e2c5 Update from knownagents.com 2026-08-25 00:43:16 +00:00
ai.robots.txt
5692a202f0 Merge pull request #271 from lyrenth/correct-aiwebindex-metadata
Correct AIWebIndex metadata
2026-08-24 14:38:04 +00:00
Glyn Normington
d8e26238a7
Merge pull request #271 from lyrenth/correct-aiwebindex-metadata
Correct AIWebIndex metadata
2026-08-24 15:37:53 +01:00
Lyrenth
763f194129 Correct AIWebIndex metadata 2026-08-24 18:27:59 +04:00
ai.robots.txt
738c80df21 Merge pull request #269 from henriquejsza/add/reflectionbot
Add Reflectionbot to robots.json
2026-08-22 09:33:25 +00:00
Glyn Normington
0a871b9e62
Merge pull request #269 from henriquejsza/add/reflectionbot
Add Reflectionbot to robots.json
2026-08-22 10:33:13 +01:00
henriquejsza
cfd62dbc3b
Add Reflectionbot to robots.json 2026-08-22 03:58:31 -03:00
Glyn Normington
2d1d8972c9
Merge pull request #267 from jamiusaliu/test/slcc-real-world-false-positives
test: pin the SLCC1/SLCC2 false positive from #208, and document whole-word matching
2026-08-20 16:05:40 +01:00
Saliu Jamiu Olamilekan
66a8329a61 docs: trim the FAQ entry per review
Applies @glyn's suggestions: less emphatic phrasing, and the SLCC1 detail
replaced by a link to issue 208 rather than restated here.

Co-authored-by: glyn <glyn@users.noreply.github.com>
2026-08-18 21:13:10 -07:00
Saliu Jamiu Olamilekan
bcc4e408e2 test: pin the SLCC1/SLCC2 false positive from #208
The non-AI user agent list covers synthetic prefix/suffix variants
(NotCursor, CursorNot). It had no real-world agent that embeds a listed
name mid-string, which is what #208 actually reported: SLCC1 and SLCC2 in
older Internet Explorer and Trident agents span the listed agent LCC.

Verified the case is load-bearing: removing the word boundaries from
list_to_pcre fails this test with
  AssertionError: <re.Match object; span=(65, 68), match='LCC'> is not None

Also document in the FAQ that agent names are matched as whole words, for
anyone consuming robots.json directly and writing their own matcher.
2026-08-18 20:28:19 -07:00
Known Agents
6a20f8dbc4 Update from knownagents.com 2026-08-18 00:41:50 +00:00
ai.robots.txt
ff88a7fcb6 Merge pull request #266 from MLuc24/add-exasearchbot
Add ExaSearchBot to robots.json
2026-08-17 16:39:01 +00:00
Glyn Normington
8e366f564b
Merge pull request #266 from MLuc24/add-exasearchbot
Add ExaSearchBot to robots.json
2026-08-17 17:38:50 +01:00
MLuc24
024c9c2d6b Add ExaSearchBot to robots.json 2026-08-17 10:30:41 +07:00
ai.robots.txt
c7c888c60e Merge pull request #263 from acook/patch-1 v1.50
add Lightpanda user agent
2026-08-17 02:39:02 +00:00
Glyn Normington
a2858c0255
Merge pull request #263 from acook/patch-1
add Lightpanda user agent
2026-08-17 03:38:51 +01:00
Glyn Normington
dd30b585f4
Merge pull request #262 from farrelldan/add-bot-ledger-related
Add Bot Ledger to Related resources
2026-08-16 04:02:44 +01:00
Anthony M. Cook
a5d17db8b8
add Lightpanda user agent
The Lightpanda browser is an AI-centric headless browser used for scraping and automation.

It is currently being used by a botnet of millions of IP addresses, including over 20k SpaceX IPs, to scrape sites such as the official Haskell Gitlab:
https://github.com/lightpanda-io/browser/issues/3156
2026-08-07 11:22:16 -05:00
Dan
abe839e4e0 docs(README): add Bot Ledger to Related resources
Free, static directory of verified AI crawlers with a one-click robots.txt/llms.txt generator, thematically aligned with this project's goal of helping site owners manage AI crawler access.
2026-08-07 08:12:59 -06:00
Glyn Normington
2f5d7ccf39
Merge pull request #260 from Ilyan321/fix/issue-257-regex-anchoring
fix(code): unanchor PCRE patterns to allow matching real user-agent headers (#257)
2026-08-05 18:18:57 +01:00
Ilyan321
4c9331ebbb Merge branch 'main' of https://github.com/ai-robots-txt/ai.robots.txt into fix/issue-257-regex-anchoring 2026-08-05 21:56:22 +05:00
ai.robots.txt
6f3054bcf0 Merge pull request #261 from fork-graveyard/main
fix lighttpd config for 1.x
2026-08-05 16:07:54 +00:00
Glyn Normington
dbd9aa78d4
Merge pull request #261 from fork-graveyard/main
fix lighttpd config for 1.x
2026-08-05 17:07:44 +01:00
girst
1177eba00a fix lighttpd config for 1.x
single quotes only allowed in 2.x. follow-up to a94b2c9.
2026-08-05 14:48:34 +02:00
Ilyan321
5658a6d726 test(code): expand false positive integration test coverage 2026-08-05 17:06:50 +05:00
Ilyan321
51a49494aa fix(code): use word boundaries \b in list_to_pcre to prevent false positive matches like NotCursor 2026-08-05 17:02:56 +05:00
Ilyan321
26510d4cd5 Merge branch 'main' of https://github.com/ai-robots-txt/ai.robots.txt into fix/issue-257-regex-anchoring 2026-08-05 16:58:22 +05:00
Ilyan321
7b1026e69c test(code): add integration test verifying regex does not match non-AI user agents 2026-08-05 16:50:35 +05:00
ai.robots.txt
5994ca0b33 Merge pull request #259 from fork-graveyard/main
do not needlessly escape hypens, make nginx matches case-sensitive, further minify regexps
2026-08-05 11:33:22 +00:00
Glyn Normington
d93db7b81d
Merge pull request #259 from fork-graveyard/main
do not needlessly escape hypens, make nginx matches case-sensitive, further minify regexps
2026-08-05 12:33:05 +01:00
Ilyan321
b06780542b fix(code): unanchor PCRE patterns to allow matching real user-agent headers (#257) 2026-08-04 23:14:02 +05:00
girst
0beb4a2666 update test files to match new output 2026-08-04 19:36:25 +02:00