Compare commits

...

459 commits

Author SHA1 Message Date
Known Agents
9ad8a47e23 Update from knownagents.com 2026-10-03 03:02:07 +00:00
Glyn Normington
e670bcafe6
Merge pull request #294 from Hitesh-XS/add-keenablebot
feat: add KeenableBot
2026-10-03 00:52:11 +01:00
SAD LIFE
341d25e239 feat: add KeenableBot 2026-10-02 10:33:53 +05:30
Known Agents
987266f3c5 Update from knownagents.com 2026-09-26 02:41:07 +00:00
Glyn Normington
5631dee581
Merge pull request #289 from dunn/faq-entries
FAQ: add entries on labor and military use
2026-09-22 19:04:41 +01:00
alexandra catalina
1ecb571700 FAQ: add references subheading, convert to unordered list 2026-09-21 10:54:24 -07:00
alexandra catalina
57ed273dc7 FAQ: add entries on labor and military use 2026-09-19 11:13:23 -07:00
Glyn Normington
6c9c1a3327 Improve wording 2026-09-19 15:47:08 +01:00
Glyn Normington
fab855244f
Merge pull request #288 from ai-robots-txt/noAI
Do not permit AI contributions
2026-09-19 15:45:13 +01:00
Glyn Normington
e4c5d5e206 Do not permit AI contributions 2026-09-19 15:43:27 +01:00
Glyn Normington
425d1a6207
Merge pull request #287 from flober81/patch-1
Add AI Discovery Radar to related resources
2026-09-17 15:25:50 +01:00
Florian Berger
9f039df105
Shorten entry: drop dated figures, spell out the file types
Updated the description of the AI Discovery Radar dataset to clarify its purpose and measurement methodology.
2026-09-17 14:01:37 +02:00
Florian Berger
8136c08d6a
Add AI Discovery Radar to related resources 2026-09-17 10:47:31 +02:00
Glyn Normington
0e111dcc24
Merge pull request #285 from SamHartleyFixes/match-agent-names-case-insensitively
Resolve agent names case-insensitively when ingesting knownagents.com
2026-09-08 06:20:04 +01:00
ai.robots.txt
197f156dd9 Merge pull request #286 from SamHartleyFixes/case-insensitive-matching
Match generated nginx/lighttpd/Caddy configs case-insensitively
2026-09-08 05:19:31 +00:00
Glyn Normington
ebc94b445f
Merge pull request #286 from SamHartleyFixes/case-insensitive-matching
Match generated nginx/lighttpd/Caddy configs case-insensitively
2026-09-08 06:19:20 +01:00
Sam Hartley
86a2e1fc67 Regenerate Caddyfile test fixture for case-insensitive match 2026-09-07 20:08:07 -07:00
Sam Hartley
3adc0cb1db Regenerate lighttpd test fixture for case-insensitive match 2026-09-07 20:08:06 -07:00
Sam Hartley
aa623355c1 Regenerate nginx test fixture for case-insensitive match 2026-09-07 20:08:06 -07:00
Sam Hartley
447bb15a22 Make the nginx/lighttpd/Caddy configs match case-insensitively 2026-09-07 20:08:05 -07:00
SamHartleyFixes
e028137f46 Resolve agent names case-insensitively when ingesting knownagents.com
robots.txt user-agent matching is case-insensitive, so two entries whose
names differ only in case are the same crawler. updated_robots_json keyed
off the scraped name, so when knownagents.com changed the capitalisation
of a name the ingest added a second entry instead of updating the first.
consolidate() also looks the name up by exact key, so the curated operator
and respect values were left behind on the old key rather than carried over.

robots.json currently carries three such pairs, and each one emits a
duplicate User-agent line in every generated file.

This resolves an incoming name against the keys already present before
using it, so an existing entry is updated in place. No generated file
changes as a result.
2026-09-07 18:08:23 -07:00
Glyn Normington
db83172548
Merge pull request #283 from SamHartleyFixes/ai-crawler-census
Add AI Crawler Census to Related
2026-09-07 18:08:07 +01:00
Sam Hartley
ee9805ef9f Add AI Crawler Census to Related 2026-09-07 08:09:25 -07:00
ai.robots.txt
2acefa38cc Merge pull request #274 from andyman01/add-qwen-ernie-doubao-mistralindex
Add four AI-lab crawlers named by Dutch news publishers
2026-09-07 03:41:26 +00:00
Glyn Normington
a595115206
Merge pull request #274 from andyman01/add-qwen-ernie-doubao-mistralindex
Add four AI-lab crawlers named by Dutch news publishers
2026-09-07 04:41:12 +01:00
andyman01
48abe91f99 Add four AI-lab crawlers named by Dutch news publishers
QwenBot, ERNIEBot, DoubaoBot and MistralAI-Index appear in the robots.txt
of Dutch news sites, measured against 83 domains in August 2026, but were
missing from this list.

MistralAI-Index is documented by the vendor at https://docs.mistral.ai/robots
and is the third of Mistral's three crawlers.

For QwenBot, ERNIEBot and DoubaoBot the evidence is different in kind and
the fields say so: the operator is identifiable from the name and the
company's own model family, but none of the three is documented on a
vendor crawler page I could find, so "respect" stays "Unclear at this
time." They are included because publishers are already naming them,
which is the same reason this list carries other undocumented agents.

Only robots.json is changed; the other files are regenerated
automatically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
2026-09-06 09:37:40 +02:00
ai.robots.txt
f76020b4e4 Merge pull request #277 from andyman01/add-kimi-mistral-training
Add KimiBot, Kimi-SearchBot and MistralAI-Training
2026-09-06 03:57:23 +00:00
Glyn Normington
b509c62249
Merge pull request #277 from andyman01/add-kimi-mistral-training
Add KimiBot, Kimi-SearchBot and MistralAI-Training
2026-09-06 04:57:15 +01:00
ai.robots.txt
da24dd28d1 Merge pull request #281 from andyman01/add-oai-adsbot
Add OAI-AdsBot
2026-09-06 03:56:13 +00:00
Glyn Normington
397d77cd74
Merge pull request #281 from andyman01/add-oai-adsbot
Add OAI-AdsBot
2026-09-06 04:56:05 +01:00
andyman01
d37bb86786 Add KimiBot, Kimi-SearchBot and MistralAI-Training
All three are documented by the vendor itself:

- KimiBot and Kimi-SearchBot: https://www.kimi.ai/policies/kimi-crawlers
  The list already has Kimi-User; Moonshot documents three crawlers.
- MistralAI-Training: https://docs.mistral.ai/robots/
  The list already has MistralAI-User. Mistral documents three; the
  third, MistralAI-Index, is proposed separately in #274.

Only robots.json is changed, per review: the other files are regenerated
automatically. Rebased on main so that Diffbot-User, merged in the
meantime, is preserved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
2026-09-05 18:58:24 +02:00
andyman01
2d7931778c Add OAI-AdsBot
Added at a maintainer's request in #277, where I raised it but did not
propose it.

OpenAI documents it at https://developers.openai.com/api/docs/bots as the
fourth of its four crawlers. Two things in that entry are worth recording
in the fields rather than leaving to the reader:

- The page does not say OAI-AdsBot obeys robots.txt, and gives no
  disallow example, so "respect" is "Unclear at this time." rather than
  the "Yes" the other three OpenAI crawlers carry.
- OpenAI states the data is not used to train foundation models, and that
  the crawler only visits pages submitted as ads. What makes it belong on
  this list is the other half of the same paragraph: content from the
  landing page is also used to "determine when it's most relevant to show
  the ad to users".

Only robots.json is changed; the other files are regenerated
automatically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
2026-09-05 15:05:22 +02:00
Glyn Normington
f44d4386fc
Merge pull request #279 from saar-twito/main
Add AI Access Checker to additional resources
2026-09-04 17:00:28 +01:00
Glyn Normington
a28a34b049
Merge pull request #280 from ai-robots-txt/instructions
Disambiguate contribution instructions
2026-09-04 16:58:06 +01:00
Glyn Normington
984382f00f Disambiguate contribution instructions
Also fix some markdown issues.
2026-09-04 16:56:43 +01:00
Glyn Normington
282ef53f61
Merge pull request #276 from andyman01/add-diffbot-user
Add Diffbot-User
2026-09-04 16:45:26 +01:00
Saar Twito
538fc8de8f
Add AI Access Checker to additional resources 2026-09-04 12:11:49 +03:00
andyman01
ed6af83530 Add Diffbot-User
Diffbot documents two separate user agents. The list currently has only
the proactive crawler; this adds the user-triggered one.

From Diffbot's own documentation:

  Diffbot-User — This is used by requests originating on behalf of a
  human user browsing a URL using Diffbot software, in response to
  their input.

Source: https://docs.diffbot.com/docs/does-crawl-respect-robotstxt

Same distinction as GPTBot / ChatGPT-User and PerplexityBot /
Perplexity-User, so it is classified as an AI Assistant fetching only
when prompted by a user.

Observed in the robots.txt of five independent publishers in four
countries: nytimes.com, bbc.co.uk, lefigaro.fr, corriere.it and
usatoday.com.

Generated files updated with code/robots.py --convert.
2026-08-29 20:34:05 +02:00
Glyn Normington
7cd4c92343
Merge pull request #272 from peppe1337/de-access-index
Add KI-Zugangsindex to Related resources
2026-08-26 23:51:59 +01:00
Ideenschmiede CEO
5e73986001 Apply maintainer's suggestion: trim entry to avoid stale figures 2026-08-26 16:02:12 +00:00
Ideenschmiede CEO
3b67c9c54e Add KI-Zugangsindex to Related resources 2026-08-25 12:15:02 +00:00
Known Agents
0037a9e2c5 Update from knownagents.com 2026-08-25 00:43:16 +00:00
ai.robots.txt
5692a202f0 Merge pull request #271 from lyrenth/correct-aiwebindex-metadata
Correct AIWebIndex metadata
2026-08-24 14:38:04 +00:00
Glyn Normington
d8e26238a7
Merge pull request #271 from lyrenth/correct-aiwebindex-metadata
Correct AIWebIndex metadata
2026-08-24 15:37:53 +01:00
Lyrenth
763f194129 Correct AIWebIndex metadata 2026-08-24 18:27:59 +04:00
ai.robots.txt
738c80df21 Merge pull request #269 from henriquejsza/add/reflectionbot
Add Reflectionbot to robots.json
2026-08-22 09:33:25 +00:00
Glyn Normington
0a871b9e62
Merge pull request #269 from henriquejsza/add/reflectionbot
Add Reflectionbot to robots.json
2026-08-22 10:33:13 +01:00
henriquejsza
cfd62dbc3b
Add Reflectionbot to robots.json 2026-08-22 03:58:31 -03:00
Glyn Normington
2d1d8972c9
Merge pull request #267 from jamiusaliu/test/slcc-real-world-false-positives
test: pin the SLCC1/SLCC2 false positive from #208, and document whole-word matching
2026-08-20 16:05:40 +01:00
Saliu Jamiu Olamilekan
66a8329a61 docs: trim the FAQ entry per review
Applies @glyn's suggestions: less emphatic phrasing, and the SLCC1 detail
replaced by a link to issue 208 rather than restated here.

Co-authored-by: glyn <glyn@users.noreply.github.com>
2026-08-18 21:13:10 -07:00
Saliu Jamiu Olamilekan
bcc4e408e2 test: pin the SLCC1/SLCC2 false positive from #208
The non-AI user agent list covers synthetic prefix/suffix variants
(NotCursor, CursorNot). It had no real-world agent that embeds a listed
name mid-string, which is what #208 actually reported: SLCC1 and SLCC2 in
older Internet Explorer and Trident agents span the listed agent LCC.

Verified the case is load-bearing: removing the word boundaries from
list_to_pcre fails this test with
  AssertionError: <re.Match object; span=(65, 68), match='LCC'> is not None

Also document in the FAQ that agent names are matched as whole words, for
anyone consuming robots.json directly and writing their own matcher.
2026-08-18 20:28:19 -07:00
Known Agents
6a20f8dbc4 Update from knownagents.com 2026-08-18 00:41:50 +00:00
ai.robots.txt
ff88a7fcb6 Merge pull request #266 from MLuc24/add-exasearchbot
Add ExaSearchBot to robots.json
2026-08-17 16:39:01 +00:00
Glyn Normington
8e366f564b
Merge pull request #266 from MLuc24/add-exasearchbot
Add ExaSearchBot to robots.json
2026-08-17 17:38:50 +01:00
MLuc24
024c9c2d6b Add ExaSearchBot to robots.json 2026-08-17 10:30:41 +07:00
ai.robots.txt
c7c888c60e Merge pull request #263 from acook/patch-1
add Lightpanda user agent
2026-08-17 02:39:02 +00:00
Glyn Normington
a2858c0255
Merge pull request #263 from acook/patch-1
add Lightpanda user agent
2026-08-17 03:38:51 +01:00
Glyn Normington
dd30b585f4
Merge pull request #262 from farrelldan/add-bot-ledger-related
Add Bot Ledger to Related resources
2026-08-16 04:02:44 +01:00
Anthony M. Cook
a5d17db8b8
add Lightpanda user agent
The Lightpanda browser is an AI-centric headless browser used for scraping and automation.

It is currently being used by a botnet of millions of IP addresses, including over 20k SpaceX IPs, to scrape sites such as the official Haskell Gitlab:
https://github.com/lightpanda-io/browser/issues/3156
2026-08-07 11:22:16 -05:00
Dan
abe839e4e0 docs(README): add Bot Ledger to Related resources
Free, static directory of verified AI crawlers with a one-click robots.txt/llms.txt generator, thematically aligned with this project's goal of helping site owners manage AI crawler access.
2026-08-07 08:12:59 -06:00
Glyn Normington
2f5d7ccf39
Merge pull request #260 from Ilyan321/fix/issue-257-regex-anchoring
fix(code): unanchor PCRE patterns to allow matching real user-agent headers (#257)
2026-08-05 18:18:57 +01:00
Ilyan321
4c9331ebbb Merge branch 'main' of https://github.com/ai-robots-txt/ai.robots.txt into fix/issue-257-regex-anchoring 2026-08-05 21:56:22 +05:00
ai.robots.txt
6f3054bcf0 Merge pull request #261 from fork-graveyard/main
fix lighttpd config for 1.x
2026-08-05 16:07:54 +00:00
Glyn Normington
dbd9aa78d4
Merge pull request #261 from fork-graveyard/main
fix lighttpd config for 1.x
2026-08-05 17:07:44 +01:00
girst
1177eba00a fix lighttpd config for 1.x
single quotes only allowed in 2.x. follow-up to a94b2c9.
2026-08-05 14:48:34 +02:00
Ilyan321
5658a6d726 test(code): expand false positive integration test coverage 2026-08-05 17:06:50 +05:00
Ilyan321
51a49494aa fix(code): use word boundaries \b in list_to_pcre to prevent false positive matches like NotCursor 2026-08-05 17:02:56 +05:00
Ilyan321
26510d4cd5 Merge branch 'main' of https://github.com/ai-robots-txt/ai.robots.txt into fix/issue-257-regex-anchoring 2026-08-05 16:58:22 +05:00
Ilyan321
7b1026e69c test(code): add integration test verifying regex does not match non-AI user agents 2026-08-05 16:50:35 +05:00
ai.robots.txt
5994ca0b33 Merge pull request #259 from fork-graveyard/main
do not needlessly escape hypens, make nginx matches case-sensitive, further minify regexps
2026-08-05 11:33:22 +00:00
Glyn Normington
d93db7b81d
Merge pull request #259 from fork-graveyard/main
do not needlessly escape hypens, make nginx matches case-sensitive, further minify regexps
2026-08-05 12:33:05 +01:00
Ilyan321
b06780542b fix(code): unanchor PCRE patterns to allow matching real user-agent headers (#257) 2026-08-04 23:14:02 +05:00
girst
0beb4a2666 update test files to match new output 2026-08-04 19:36:25 +02:00
girst
a94b2c93e5 further minify regexps
- nginx and lighttpd allow single quoted strings, so use python's
  built-in methods to escape quotes
- only apache requires parentheses on the outside (and only for obscure
  reasons)
- caddy expects double quoted strings. at least escape inner quotes
  directly (re.escape does not do this since python 3.7), should any
  appear
2026-08-04 14:48:03 +02:00
girst
3bd200ba32 make nginx matches case-sensitive
fixes #225.
2026-08-04 14:18:17 +02:00
girst
c4b366740e do not needlessly escape hypens
partially fixes #225.
2026-08-04 14:13:52 +02:00
Glyn Normington
27389c0b3d
Merge pull request #256 from guest20/patch-1
robots.py: duplicate parser = argparse.ArgumentParser(...)
2026-08-04 12:05:31 +01:00
Glyn Normington
940bf7f7e4
Merge pull request #255 from INXPRNCD/data/fill-unclear-6-entries
Fill "Unclear at this time." from primary operator docs for 6 entries
2026-08-04 11:59:32 +01:00
Glyn Normington
1e031f41d8
Merge pull request #254 from INXPRNCD/fix/ci-commit-message-injection
Pass commit messages to bash via env, not template interpolation
2026-08-04 11:57:47 +01:00
guest20
265ea03066
robots.py: duplciate parser = argparse.ArgumentParser(...)
Two is just greedy
2026-08-04 10:24:30 +02:00
Özden und Julia
2eb9315743 Fill "Unclear at this time." from primary operator docs for 6 entries
Re-submission of #248, which was merged then reverted in 80c19fc. The revert was
caused by main.yml's commit-message handling, not by this data — that is fixed
separately in #254.

Six entries move from "Unclear at this time." to the operator's own documented
value, each citing the primary source inline:

  Applebot              Yes   support.apple.com/en-us/119829#retrieval
  Meta-ExternalAgent    Yes   developers.facebook.com/docs/sharing/webmasters/web-crawlers/
  Meta-ExternalFetcher  No    (same Meta page — user-initiated fetch, documented
  meta-externalfetcher  No     as not checking robots.txt)
  DuckAssistBot         Yes   duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/
  Google-Agent          Yes   developers.google.com/search/docs/crawling-indexing/google-common-crawlers#google-agent

table-of-bot-metrics.md is regenerated and included, so `robots.py --convert` has
nothing left to write and the workflow's commit step is skipped entirely.

Verified: code/tests.py 13/13 · robots.py --convert exits 0 and is idempotent ·
robots.txt, .htaccess, the nginx/lighttpd/haproxy configs and the Caddyfile are
byte-identical · the JSON diff touches only the six `respect` fields, nothing else.
2026-08-04 09:11:42 +02:00
Özden und Julia
4021a10239 Pass commit messages to bash via env, not template interpolation
`git commit -m "${{ github.event.head_commit.message }}"` splices arbitrary text
straight into a double-quoted bash string. A message containing a double quote
closes the string early and the remainder is re-parsed as shell words.

This is what broke main after #248 merged. That PR's title contained
"Unclear at this time." (with quotes), so the merge commit message did too, and
the line bash actually ran was:

    git commit -m "Fill "Unclear at this time." from primary operator docs ..."

which git received as:

    -m "Fill Unclear"  at  this  "time. from primary operator docs ..."

Hence `error: pathspec 'at' did not match any file(s) known to git`, a failed
run, and the revert in 80c19fc. The data in that PR was fine — robots.py
--convert exits 0 against it and code/tests.py passes 13/13.

The same interpolation is also a script-injection vector, which is the more
important reason to change it: a PR title is attacker-controlled, and

    chore: tidy"; <any command>; echo "

executes that command on the runner with the workflow's token. `inputs.message`
has the same shape in the `if [ -n ... ]` test and its own `git commit -m`, so
all three are moved.

Passing through `env:` and quoting the shell variable is GitHub's documented
recommendation for untrusted values. The variable is expanded by bash after
parsing, so quotes, newlines and `$(...)` stay literal text.

Verified: code/tests.py 13/13 · python code/robots.py --convert exits 0 and
leaves robots.txt, table-of-bot-metrics.md and every server-config output
byte-identical · reproduced both the parse failure and the injection locally
against the old form, and confirmed the env form commits the same message
verbatim, quotes included.
2026-08-04 09:09:10 +02:00
ai.robots.txt
24c010e668 Merge pull request #251 from glyn/delete-IbouBot
Delete IbouBot
2026-07-31 15:01:58 +00:00
Glyn Normington
1be8fa6220
Merge pull request #251 from glyn/delete-IbouBot
Delete IbouBot
2026-07-31 16:01:41 +01:00
Glyn Normington
69a87566de Delete IbouBot
Fixes https://github.com/ai-robots-txt/ai.robots.txt/issues/250
2026-07-31 13:04:31 +01:00
Glyn Normington
042d00a5ab
Merge pull request #249 from ai-robots-txt/test-ci
Empty commit to test CI
2026-07-31 03:35:15 +01:00
Glyn Normington
c96a9916cc Empty commit to test CI 2026-07-31 03:34:01 +01:00
Glyn Normington
80c19fcad5 Revert "Fill operator and respect from primary docs for 6 entries"
This reverts commit e62a95bb8a.

Reverted because of CI failure noted in PR.
2026-07-31 03:30:04 +01:00
Glyn Normington
bbe0579a23
Merge pull request #248 from INXPRNCD/fill-unclear-operator-and-respect-from-primary-docs
Fill "Unclear at this time." from primary operator docs for 6 entries
2026-07-31 03:20:58 +01:00
Known Agents
2155906baa Update from knownagents.com 2026-07-31 01:57:31 +00:00
Erdinc Özden
e62a95bb8a Fill operator and respect from primary docs for 6 entries
Applebot, Meta-ExternalAgent, Meta-ExternalFetcher, meta-externalfetcher,
DuckAssistBot and Google-Agent carried 'Unclear at this time.' in both
operator and respect, but each operator documents the behaviour on its
own site. Values are linked to the primary source inline, matching the
style already used by Applebot-Extended and meta-externalagent.

robots.txt output is unchanged — no keys added or renamed.
2026-07-30 22:05:49 +02:00
Glyn Normington
3db35be22d
Merge pull request #241 from iamdhrv/fix/exact-spider-user-agent
Anchor generic Spider user-agent match
2026-07-30 12:15:55 +01:00
Dhruv Maniya
49669178ef fix: configure exact user-agent matching
Signed-off-by: Dhruv Maniya <dhruvmaniya1998@gmail.com>
2026-07-29 21:47:42 +05:30
Dhruv Maniya
af31418609 Anchor the generic Spider user agent
Signed-off-by: Dhruv Maniya <dhruvmaniya1998@gmail.com>
2026-07-15 13:04:57 +05:30
Known Agents
a0fed45f0e Update from knownagents.com 2026-07-10 02:03:20 +00:00
ai.robots.txt
2a83168bbb Merge pull request #240 from bobmatyas/add/tabstack
Add: Mozilla-Tabstack
2026-07-09 15:01:12 +00:00
Glyn Normington
51f7fd5e3a
Merge pull request #240 from bobmatyas/add/tabstack
Add: Mozilla-Tabstack
2026-07-09 16:00:59 +01:00
Known Agents
a71b27a232 Update from knownagents.com 2026-07-09 02:04:28 +00:00
bob
c793fa0a8c Add: Mozilla-Tabstack. Closes #226 2026-07-08 21:16:52 -04:00
Glyn Normington
f030cbe956
Merge pull request #238 from ghking/rename-dark-visitors-to-known-agents
Renaming Dark Visitors to Known Agents
2026-07-09 01:41:52 +01:00
Glyn Normington
12a15a3a28
Update code/robots.py
For consistency.
2026-07-09 01:40:00 +01:00
Glyn Normington
5065ca6055
Update .github/workflows/ai_robots_update.yml
For consistency.
2026-07-09 01:39:50 +01:00
Glyn Normington
0b021b9e40
Update code/robots.py
For consistency.
2026-07-09 01:39:39 +01:00
Gavin King
dad7db5168 Fixed broken URLs 2026-07-08 20:38:48 -04:00
Gavin King
e1bf273b16 Update robots.py 2026-07-08 20:26:48 -04:00
Gavin King
a35e6f9a09 Update ai_robots_update.yml 2026-07-08 20:23:53 -04:00
Gavin King
c43c5afd07 More name changes 2026-07-08 15:43:07 -04:00
Gavin King
0c5de5b90b Updating the name 2026-07-08 15:30:25 -04:00
dark-visitors
a394130efd Update from Dark Visitors 2026-06-26 02:32:55 +00:00
Glyn Normington
ee4ae068c2
Merge pull request #236 from guest20/patch-2
robots.py: fix darkvisiters descriptions and links
2026-06-23 11:15:59 +01:00
Glyn Normington
82df5f7659
Apply suggestions from code review
Co-authored-by: Glyn Normington <work@underlap.org>
2026-06-23 11:13:17 +01:00
guest20
ac207f9947
tests.py: ) 2026-06-23 04:17:14 +02:00
guest20
ffad9293c4
tests.py: consolidate 2026-06-23 04:13:17 +02:00
guest20
80357baf9c
robots.py: pass everything to consolodate 2026-06-21 23:51:32 +02:00
guest20
7dbe93295e
robots.py: pull consolodate out 2026-06-21 23:47:11 +02:00
guest20
3f965b292d
robots.py: keep new, non-default values
Co-authored-by: Glyn Normington <work@underlap.org>
2026-06-20 18:29:47 +02:00
dark-visitors
5cb3ec765f Update from Dark Visitors 2026-06-17 11:21:12 +00:00
guest20
25415e342c
robots.py: div class=description 2026-06-17 13:20:29 +02:00
dark-visitors
118bb772c0 Update from Dark Visitors 2026-06-17 11:01:38 +00:00
guest20
8dd8e6c4e4
robots.py: possibly ruin the whole purpose of consolodate
Replace an old non-default value with a new non-default value, so updates happen.
2026-06-17 13:00:15 +02:00
guest20
78fd7b49d8
\robots.py 2026-06-17 12:23:10 +02:00
guest20
101d7ac0ad
robots.py: punked by auto-) again 2026-06-17 12:10:05 +02:00
guest20
f2147fd81c
robots.py: fix darkvistors link 2026-06-17 12:06:50 +02:00
guest20
bd3ee4702c
robots.py: agents/agents -> agents 2026-06-17 11:45:55 +02:00
dark-visitors
3091ad0a23 Update from Dark Visitors 2026-06-16 02:52:42 +00:00
dark-visitors
f420408eee Update from Dark Visitors 2026-06-04 02:48:49 +00:00
ai.robots.txt
000552c2df Merge pull request #232 from justinribeiro/main
add: AgentTimes bot
2026-06-03 08:56:38 +00:00
Glyn Normington
ef50defc1a
Merge pull request #232 from justinribeiro/main
add: AgentTimes bot
2026-06-03 09:56:26 +01:00
Justin Ribeiro
7673437733
add: AgentTimes bot 2026-06-02 11:26:06 -07:00
dark-visitors
ee3a96ee5c Update from Dark Visitors 2026-05-22 02:33:30 +00:00
Glyn Normington
1fbf7a06ac
Merge pull request #231 from cfcryptotrader-rgb/cfcryptotrader-sync-generated-robots
Keep generated robot files synced after Dark Visitors update
2026-05-13 07:44:48 +01:00
cfcryptotrader-rgb
8d5a081a16 Keep generated robot files in dark visitors update 2026-05-13 00:44:06 -03:00
dark-visitors
aea7db9e34 Update from Dark Visitors 2026-05-13 02:24:20 +00:00
Glyn Normington
767fcc6f15
Merge pull request #229 from axeleroy/main
Add missing bot categories
2026-05-12 04:33:19 +01:00
Axel Leroy
6adc1f7b53
Add missing bot categories 2026-05-11 15:42:56 +02:00
dark-visitors
851ce068cf Update from Dark Visitors 2026-05-03 02:01:44 +00:00
ai.robots.txt
8b0bc322f4 Merge pull request #227 from gsauthof/naget-news
Add NagetBot and newsai user-agents
2026-05-02 15:05:09 +00:00
Glyn Normington
7244a7e870
Merge pull request #227 from gsauthof/naget-news
Add NagetBot and newsai user-agents
2026-05-02 16:04:59 +01:00
Georg Sauthoff
a9f686064f Add NagetBot and newsai user-agents
User-agent header values seen in the wild over the last days:

    Mozilla/5.0 (compatible; NagetBot/1.0; +https://naget.ai/bot)
    Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 newsai/1.0 Safari/537.36
2026-05-02 13:04:51 +02:00
ai.robots.txt
198653b59a Update from Dark Visitors 2026-03-27 01:26:26 +00:00
dark-visitors
ad47cafc97 Update from Dark Visitors 2026-03-26 01:26:23 +00:00
dark-visitors
2d2ec4ce8d Update from Dark Visitors 2026-03-25 01:20:42 +00:00
ai.robots.txt
86d582b11c Update from Dark Visitors 2026-03-07 01:12:48 +00:00
dark-visitors
3de8865f00 Update from Dark Visitors 2026-03-06 01:21:01 +00:00
Glyn Normington
243ec6b67d
Merge pull request #216 from flymarq/lighttpd
Lighttpd
2026-02-17 17:28:34 +00:00
Flynn Marquardt
793b0454f2 Use adverb. 2026-02-15 11:20:34 +01:00
Flynn Marquardt
790bff1817 Add reference file for lighttpd test. 2026-02-15 11:18:35 +01:00
Flynn Marquardt
742f9b9b5e Add lighttpd test. 2026-02-15 11:17:54 +01:00
ai.robots.txt
48bfefbc5f Update from Dark Visitors 2026-02-14 01:16:06 +00:00
Flynn Marquardt
5f0dd97ccb Add lighttpd sample file. 2026-02-13 19:23:19 +01:00
Flynn Marquardt
55d404c115 Fix typos. 2026-02-13 19:21:31 +01:00
Flynn Marquardt
deae8eed2d Update with lighttpd instructions. 2026-02-13 19:20:12 +01:00
Flynn Marquardt
70e83e6056 Add generation and output of lighttpd configuration fragment. 2026-02-13 19:18:46 +01:00
dark-visitors
a762015e58 Update from Dark Visitors 2026-02-13 01:23:17 +00:00
ai.robots.txt
aa8519ec10 Update from Dark Visitors 2025-12-21 01:07:06 +00:00
dark-visitors
83485effdb Update from Dark Visitors 2025-12-20 00:58:49 +00:00
ai.robots.txt
8b8bf9da5d Update from Dark Visitors 2025-12-06 00:58:16 +00:00
dark-visitors
f1c752ef12 Update from Dark Visitors 2025-12-05 01:00:44 +00:00
Adam Newbold
51afa7113a
Update ai_robots_update.yml with rebase command to fix scheduled run 2025-12-03 20:15:10 -05:00
ai.robots.txt
7598d77e4a Update from Dark Visitors 2025-12-04 01:10:02 +00:00
dark-visitors
45b071b29f Update from Dark Visitors 2025-12-04 01:00:27 +00:00
dark-visitors
f61b3496f7 Update from Dark Visitors 2025-12-03 01:00:32 +00:00
dark-visitors
8363d4fdd4 Update from Dark Visitors 2025-12-02 01:25:24 +00:00
Adam Newbold
2ccd443581
Update ai_robots_update.yml with workflow_dispatch
Adding workflow_dispatch to enable manual triggers of this schedule job (for testing)
2025-12-01 20:24:41 -05:00
Adam Newbold
6d75f3c1c9
Update robots.py to address error on line 57
Attempting to work around an error that prevents parsing the Dark Visitors site
2025-12-01 20:18:29 -05:00
Glyn Normington
56010ef913
Merge pull request #205 from fiskhandlarn/fix/editorconfig
Fix/editorconfig
2025-11-29 10:02:08 +00:00
ai.robots.txt
3fadc88a23 Merge pull request #206 from newbold/main
Adding LAIONDownloader
2025-11-29 10:00:52 +00:00
Glyn Normington
47c077a8ef
Merge pull request #206 from newbold/main
Adding LAIONDownloader
2025-11-29 10:00:42 +00:00
Adam Newbold
f5d7ccb243
Fixed invalid JSON 2025-11-28 14:10:38 -05:00
Adam Newbold
30d719a09a
Adding LAIONDownloader 2025-11-28 14:08:18 -05:00
fiskhandlarn
05bbdebeaa feat: disallow final newline for files generated by python
if any of these files have ending newlines the tests will fail
2025-11-28 10:47:22 +01:00
fiskhandlarn
c6ce9329a1 fix: ensure whitespace as defined in .editorconfig 2025-11-28 10:46:22 +01:00
ai.robots.txt
10d5ae2870 Merge pull request #200 from glyn/deepseek
Clarify that DeepSeekBot does not respect robots.txt
2025-11-27 14:00:50 +00:00
Glyn Normington
4467002298
Merge pull request #200 from glyn/deepseek
Clarify that DeepSeekBot does not respect robots.txt
2025-11-27 14:00:40 +00:00
Glyn Normington
4a159a818f
Merge pull request #199 from glyn/editorconfig
Standardise editor options
2025-11-27 13:54:03 +00:00
Glyn Normington
3d6b33a71a
Merge pull request #202 from glyn/formatting
Tidy README
2025-11-27 12:42:31 +00:00
Glyn Normington
c26c8c0911 Tidy README 2025-11-27 12:41:37 +00:00
Glyn Normington
91959fe791
Merge pull request #201 from Anshita-18H/add-requirements-file
Add requirements.txt with project dependencies
2025-11-27 12:38:28 +00:00
Glyn Normington
b75163e796
Ensure Python3 is used 2025-11-27 12:38:12 +00:00
Glyn Normington
f46754d280
Order deps 2025-11-27 12:37:35 +00:00
Glyn Normington
c3f2fe758e
Whitespace 2025-11-27 12:37:21 +00:00
Glyn Normington
7521c3af50
Fix link 2025-11-27 12:37:07 +00:00
Anshita-18H
4302fd1aca Improve contributing section and fix formatting in README 2025-11-27 17:56:47 +05:30
Anshita-18H
9ca5033927 Deduplicate requirements.txt and add installation instructions to README 2025-11-27 17:44:47 +05:30
Anshita-18H
9cc3dbc05f Add requirements.txt with project dependencies 2025-11-27 17:06:59 +05:30
Glyn Normington
c322d7852d
Merge pull request #130 from maxheadroom/main
Create traefik-manual-setup.md
2025-11-26 12:08:11 +00:00
Glyn Normington
35588b2ddb
Typo. 2025-11-26 12:07:39 +00:00
Glyn Normington
ab24e41106 Clarify that DeepSeekBot does not respect robots.txt
Fixes https://github.com/ai-robots-txt/ai.robots.txt/issues/198
2025-11-26 12:02:42 +00:00
Glyn Normington
b681c0c0d8 Standardise editor options
This was motivated by @fiskhandlarn's comment:
https://github.com/ai-robots-txt/ai.robots.txt/pull/195#issuecomment-3576331322
2025-11-26 11:54:10 +00:00
Glyn Normington
4e7e28335f
Merge pull request #195 from fiskhandlarn/patch-2
feat: allow robots access to `/robots.txt` in nginx
2025-11-26 11:50:15 +00:00
fiskhandlarn
a6cf6b204b test: update test nginx conf 2025-11-25 16:47:04 +01:00
fiskhandlarn
2679fcad34 feat: update nginx generator 2025-11-25 16:46:31 +01:00
fiskhandlarn
ef8eda4fe6 chore: normalize quote style 2025-11-25 16:39:35 +01:00
Adam Newbold
a29102f0fc
Merge pull request #196 from glyn/releasing
Document how to ship a new release
2025-11-25 10:34:25 -05:00
fiskhandlarn
0b3266b35f feat: allow robots access to /robots.txt in nginx 2025-11-25 16:22:14 +01:00
Glyn Normington
6225e3e98e Document how to ship a new release 2025-11-23 04:01:43 +00:00
ai.robots.txt
be4d74412c Merge pull request #194 from glyn/193-kendra
delete extraneous hyphen
2025-11-21 16:32:11 +00:00
Cory Dransfeldt
729be4693a
Merge pull request #194 from glyn/193-kendra
delete extraneous hyphen
2025-11-21 08:32:01 -08:00
Glyn Normington
efb4d260da delete extraneous hyphen
Fixes https://github.com/ai-robots-txt/ai.robots.txt/issues/193
2025-11-20 10:23:48 +00:00
ai.robots.txt
e2726ac160 Merge pull request #192 from ai-robots-txt/cdransf/notebooklm-klaviyo
chore: adds NotebookLM and KlaviyoAIBot agents
2025-11-14 21:01:25 +00:00
Glyn Normington
663d030f96
Merge pull request #192 from ai-robots-txt/cdransf/notebooklm-klaviyo
chore: adds NotebookLM and KlaviyoAIBot agents
2025-11-14 21:01:13 +00:00
Cory Dransfeldt
28b45ea08d
chore: adds NotebookLM and KlaviyoAIBot agents 2025-11-14 11:03:03 -08:00
ai.robots.txt
443dd27527 Merge pull request #189 from ai-robots-txt/cdransf/atlassian-amazon-bots
chore(robots.json): add AmazonBuyForMe and atlassian-bot
2025-11-05 23:50:59 +00:00
Glyn Normington
60b6a0829d
Merge pull request #189 from ai-robots-txt/cdransf/atlassian-amazon-bots
chore(robots.json): add AmazonBuyForMe and atlassian-bot
2025-11-05 23:50:48 +00:00
Cory Dransfeldt
00bf2b0e13
chore(robots.json): add AmazonBuyForMe and atlassian-bot 2025-11-05 14:02:37 -08:00
ai.robots.txt
e87eb706e3 Merge pull request #188 from ai-robots-txt/cdransf/buddybot
chore: adds BuddyBot
2025-11-04 05:16:28 +00:00
Glyn Normington
3d41350256
Merge pull request #188 from ai-robots-txt/cdransf/buddybot
chore: adds BuddyBot
2025-11-04 05:16:15 +00:00
Cory Dransfeldt
808451055c
chore: adds BuddyBot 2025-11-03 13:39:07 -08:00
ai.robots.txt
5cad0ee389 Merge pull request #187 from ai-robots-txt/cdransf/add-Linguee-Bot
chore: adds Linguee Bot
2025-10-24 18:05:01 +00:00
Glyn Normington
d55c9980cd
Merge pull request #187 from ai-robots-txt/cdransf/add-Linguee-Bot
chore: adds Linguee Bot
2025-10-24 19:04:42 +01:00
Cory Dransfeldt
192b0a2eef
chore: adds Linguee Bot 2025-10-24 10:35:04 -07:00
dark-visitors
97e19445ce Update from Dark Visitors 2025-10-24 00:53:30 +00:00
ai.robots.txt
0bc2361be8 Merge pull request #186 from ai-robots-txt/third-party
Use third party in text, rather than first
2025-10-21 15:23:46 +00:00
Cory Dransfeldt
511d8c955d
Merge pull request #186 from ai-robots-txt/third-party
Use third party in text, rather than first
2025-10-21 08:23:28 -07:00
Glyn Normington
b89a9eae6a Use third party in text, rather than first 2025-10-21 01:11:32 +01:00
ai.robots.txt
646ab08e15 Merge pull request #185 from ai-robots-txt/cdransf/add-IbouBot
chore: adds IbouBot
2025-10-21 00:09:24 +00:00
Glyn Normington
5c3da1c1af
Merge pull request #185 from ai-robots-txt/cdransf/add-IbouBot
chore: adds IbouBot
2025-10-21 01:09:06 +01:00
Glyn Normington
d22b8dfd7b
Description is third party, not first 2025-10-21 01:08:42 +01:00
Cory Dransfeldt
19c1d346c3
chore: adds IbouBot 2025-10-20 09:28:31 -07:00
ai.robots.txt
2fa0e9119c Merge pull request #183 from ai-robots-txt/cdransf/bot-additions-cloudflare
chore: add amazon-kendra-, Anomura, Cloudflare-AutoRAG and Bravebot
2025-10-20 14:34:43 +00:00
Glyn Normington
0874a92503
Merge pull request #183 from ai-robots-txt/cdransf/bot-additions-cloudflare
chore: add amazon-kendra-, Anomura, Cloudflare-AutoRAG and Bravebot
2025-10-20 15:34:30 +01:00
Cory Dransfeldt
28d2d09633
chore: add amazon-kendra-, Anomura, Cloudflare-AutoRAG and Bravebot 2025-10-19 15:57:28 -07:00
ai.robots.txt
260f5029fe Update from Dark Visitors 2025-10-16 00:56:33 +00:00
dark-visitors
91bf905fa9 Update from Dark Visitors 2025-10-15 00:56:30 +00:00
dark-visitors
56d03d46fb Update from Dark Visitors 2025-09-26 00:54:30 +00:00
ai.robots.txt
38d60b928c Merge pull request #179 from ai-robots-txt/deepseekbot
chore(robots.json): add DeepSeekBot
2025-09-25 10:17:38 +00:00
Glyn Normington
e2266bbc1d
Merge pull request #179 from ai-robots-txt/deepseekbot
chore(robots.json): add DeepSeekBot
2025-09-25 11:17:23 +01:00
Cory Dransfeldt
bf347bdf91
chore(robots.json): add DeepSeekBot 2025-09-24 09:42:01 -07:00
nisbet-hubbard
c6e7d69dd5
Update README.md 2025-09-24 09:39:21 -07:00
ai.robots.txt
8906f6b447 Update from Dark Visitors 2025-09-12 00:53:06 +00:00
dark-visitors
2fd93029ca Update from Dark Visitors 2025-09-11 00:55:06 +00:00
ai.robots.txt
b6338ddc73 Merge pull request #176 from emersion/TerraCotta
Add TerraCotta
2025-09-10 14:11:16 +00:00
Cory Dransfeldt
7ffbf33baf
Merge pull request #176 from emersion/TerraCotta
Add TerraCotta
2025-09-10 07:11:04 -07:00
Simon Ser
50870ba911 Add TerraCotta
Example server log:

    X.X.X.X - - [10/Sep/2025:00:06:17 +0000] "GET /archives/dri-devel/2023-November/430875.html HTTP/1.1" 200 7388 "-" "TerraCotta https://github.com/CeramicTeam/CeramicTerracotta"
2025-09-10 12:04:08 +02:00
dark-visitors
dd391bf960 Update from Dark Visitors 2025-09-10 00:53:56 +00:00
ai.robots.txt
229d1b4dbc Merge pull request #174 from karolyi/master
Update Brightbot operator and details; add meta-webindexer entry
2025-09-09 02:45:32 +00:00
Glyn Normington
4d506ca322
Merge pull request #174 from karolyi/master
Update Brightbot operator and details; add meta-webindexer entry
2025-09-09 03:45:23 +01:00
László Károlyi
ec508ab434
Update Brightbot operator and details; add meta-webindexer entry
- Update Brightbot operator to https://brightdata.com/brightbot.
- Change Brightbot frequency to "At least one per minute."
- Expand Brightbot description with disguise tactics link.
- Add new entry for meta-webindexer under Meta operator.
- Set meta-webindexer respect to "Unclear at this time."
- Define meta-webindexer function as "AI Assistants."
- Set meta-webindexer frequency to "Unhinged, more than 1 per second."
- Include meta-webindexer description on improving Meta AI search.
2025-09-08 21:42:06 +02:00
dark-visitors
0ed29412c9 Update from Dark Visitors 2025-08-28 00:55:47 +00:00
ai.robots.txt
2ad1c3e831 Merge pull request #168 from ai-robots-txt/google-firebase-shapbot
chore: add Google-Firebase and ShapBot
2025-08-27 10:13:17 +00:00
Glyn Normington
1677278c5a
Merge pull request #168 from ai-robots-txt/google-firebase-shapbot
chore: add Google-Firebase and ShapBot
2025-08-27 11:13:04 +01:00
Cory Dransfeldt
1a8edfa84a
chore: add Google-Firebase and ShapBot 2025-08-26 19:39:42 -07:00
dark-visitors
cf073d49f2 Update from Dark Visitors 2025-08-15 01:01:42 +00:00
ai.robots.txt
6d3f3e1712 Merge pull request #167 from ai-robots-txt/OpenAI-bot
chore: adds OpenAI but user agent
2025-08-14 01:45:05 +00:00
Glyn Normington
784b8440a5
Merge pull request #167 from ai-robots-txt/OpenAI-bot
chore: adds OpenAI but user agent
2025-08-14 02:44:53 +01:00
Cory Dransfeldt
0e687a5b58
chore: adds OpenAI but user agent 2025-08-13 09:01:55 -07:00
ai.robots.txt
ff9fc26404 Update from Dark Visitors 2025-08-01 01:12:53 +00:00
dark-visitors
146a229662 Update from Dark Visitors 2025-07-31 01:05:21 +00:00
ai.robots.txt
64f9d6ce9c Update from Dark Visitors 2025-07-30 01:05:21 +00:00
dark-visitors
085dd1071e Update from Dark Visitors 2025-07-29 01:11:26 +00:00
ai.robots.txt
9565c11d4c Merge pull request #166 from ai-robots-txt/summaly-bot
fix: remove summaly
2025-07-28 19:19:15 +00:00
Glyn Normington
8869442615
Merge pull request #166 from ai-robots-txt/summaly-bot
fix: remove summaly
2025-07-28 20:19:06 +01:00
ai.robots.txt
27420d6fed Merge pull request #164 from GlitzSmarter/main
Adding YaK
2025-07-28 18:29:57 +00:00
Glyn Normington
44e58a5ece
Merge pull request #164 from GlitzSmarter/main
Adding YaK
2025-07-28 19:29:48 +01:00
Glyn Normington
12c5368e04
Fix syntax 2025-07-28 19:29:19 +01:00
Glitz Smarter
c15065544a
Rewriting function in YaK
Rewrote the function to make it clear the quote is according to the page on the Clearwater website.
2025-07-28 12:54:25 -04:00
Cory Dransfeldt
9171625db6
fix: remove summaly 2025-07-28 08:56:21 -07:00
Gregory
8b188d0612 Added YaK user-agent to list of robots
Added Meltwater's AI to list of bots that use AI and scrape websites.
2025-07-26 23:50:11 -04:00
ai.robots.txt
47769c8429 Update from Dark Visitors 2025-07-17 01:04:20 +00:00
dark-visitors
9843c31303 Update from Dark Visitors 2025-07-16 01:03:43 +00:00
ai.robots.txt
567e94a6cf Merge pull request #160 from sdavids/users/sdavids/feat/htaccess
Simplify htaccess rewrite rule
2025-07-08 16:00:16 +00:00
Glyn Normington
30c4c037d9
Merge pull request #160 from sdavids/users/sdavids/feat/htaccess
Simplify htaccess rewrite rule
2025-07-08 16:59:52 +01:00
Sebastian Davids
c7407e721d
Simplify htaccess rewrite rule
https://httpd.apache.org/docs/2.4/rewrite/flags.html#flag_f

fix #159

Signed-off-by: Sebastian Davids <sdavids@gmx.de>
2025-07-08 12:33:24 +02:00
dark-visitors
f6f504012f Update from Dark Visitors 2025-07-08 01:01:25 +00:00
ai.robots.txt
63f1f56307 Merge pull request #158 from ai-robots-txt/thinkbot-summaly
chore(robots.json): add SummalyBot + Thinkbot
2025-07-07 16:55:31 +00:00
Glyn Normington
91a5b3c995
Merge pull request #158 from ai-robots-txt/thinkbot-summaly
chore(robots.json): add SummalyBot + Thinkbot
2025-07-07 17:55:22 +01:00
Cory Dransfeldt
c5dd4e98b6
chore(robots.json): add SummalyBot + Thinkbot 2025-07-07 09:49:26 -07:00
dark-visitors
70ff0c9fbd Update from Dark Visitors 2025-07-03 01:00:49 +00:00
ai.robots.txt
20a242d390 Merge pull request #156 from nisbet-hubbard/patch-9
Update funtion and description for SemrushBot
2025-07-02 15:58:19 +00:00
Cory Dransfeldt
91ad302d68
Merge pull request #156 from nisbet-hubbard/patch-9
Update funtion and description for SemrushBot
2025-07-02 08:58:08 -07:00
dark-visitors
79bc22523f Update from Dark Visitors 2025-07-02 01:01:05 +00:00
ai.robots.txt
2bfcee794d Update from Dark Visitors 2025-06-29 01:07:47 +00:00
dark-visitors
7eb2099d3f Update from Dark Visitors 2025-06-28 00:58:57 +00:00
nisbet-hubbard
accc48f327
Update funtion and description for SemrushBot 2025-06-21 08:48:58 +08:00
dark-visitors
4ed17b8e4a Update from Dark Visitors 2025-06-17 01:00:21 +00:00
ai.robots.txt
5326c202b5 Merge pull request #154 from paulrudy/main
re-add facebookexternalhit
2025-06-16 15:12:42 +00:00
Cory Dransfeldt
a31ae1e6d0
Merge pull request #154 from paulrudy/main
re-add facebookexternalhit
2025-06-16 08:12:31 -07:00
paulrudy
7535893aec re-add facebookexternalhit 2025-06-15 16:49:07 -07:00
ai.robots.txt
eb05f2f527 Merge pull request #153 from sergiospagnuolo/Poseidon
Update robots.json with new crawler
2025-06-14 14:04:03 +00:00
Cory Dransfeldt
26a46c409d
Merge pull request #153 from sergiospagnuolo/Poseidon
Update robots.json with new crawler
2025-06-14 07:03:52 -07:00
dark-visitors
2b68568ac2 Update from Dark Visitors 2025-06-14 00:58:11 +00:00
Sérgio Spagnuolo
b05f2fee00
Update robots.json with new crawler
Update with Poseidon Research Crawler as found in nytimes.com/robots.txt
2025-06-13 17:15:13 -03:00
ai.robots.txt
e53d81c66d Merge pull request #152 from ai-robots-txt/MyCentralAIScraperBot
chore(robots.json): adds MyCentralAIScraperBot
2025-06-13 09:28:41 +00:00
Glyn Normington
20e327e74e
Merge pull request #152 from ai-robots-txt/MyCentralAIScraperBot
chore(robots.json): adds MyCentralAIScraperBot
2025-06-13 10:28:32 +01:00
Glyn Normington
8f17718e76
Fix typo 2025-06-13 10:28:12 +01:00
Cory Dransfeldt
d760f9216f
chore(robots.json): adds MyCentralAIScraperBot 2025-06-12 13:08:29 -07:00
ai.robots.txt
842e2256e8 Merge pull request #150 from ai-robots-txt/semrush-bots
chore(robots.json): adds additional SemrushBot user agents
2025-06-12 07:12:00 +00:00
Glyn Normington
229ea20426
Merge pull request #150 from ai-robots-txt/semrush-bots
chore(robots.json): adds additional SemrushBot user agents
2025-06-12 08:11:51 +01:00
Cory Dransfeldt
14d68f05ba
chore(robots.json): adds additional SemrushBot user agents 2025-06-11 13:50:53 -07:00
dark-visitors
cf598b6b71 Update from Dark Visitors 2025-06-10 01:00:37 +00:00
ai.robots.txt
3759a6bf14 chore(robots.json): adds EchoboxBot (#148) 2025-06-09 15:44:36 +00:00
Cory Dransfeldt
7867c3e26c
chore(robots.json): adds EchoboxBot (#148) 2025-06-09 16:44:25 +01:00
dark-visitors
e21f6ae1b6 Update from Dark Visitors 2025-06-06 00:59:25 +00:00
ai.robots.txt
ac7ed17e71 Merge pull request #145 from ai-robots-txt/aws-bedrockbot
chore(robots.json): adds bedrockbot
2025-06-05 16:51:17 +00:00
Glyn Normington
81747e6772
Merge pull request #145 from ai-robots-txt/aws-bedrockbot
chore(robots.json): adds bedrockbot
2025-06-05 17:51:03 +01:00
Cory Dransfeldt
528d77bf07
chore(robots.json): adds bedrockbot 2025-06-05 09:14:23 -07:00
dark-visitors
77393df5aa Update from Dark Visitors 2025-06-05 00:59:28 +00:00
ai.robots.txt
75ea75a95b Merge pull request #143 from ai-robots-txt/panscient
chore(robots.json): adds Panscient
2025-06-04 18:04:06 +00:00
Glyn Normington
2fca1ddcf1
Merge pull request #143 from ai-robots-txt/panscient
chore(robots.json): adds Panscient
2025-06-04 19:03:53 +01:00
ai.robots.txt
9c28c63a0c Merge pull request #142 from ai-robots-txt/quillbot
chore(robots.json): adds Quillbot
2025-06-04 17:54:57 +00:00
Cory Dransfeldt
395c013eea
Merge pull request #142 from ai-robots-txt/quillbot
chore(robots.json): adds Quillbot
2025-06-04 10:54:46 -07:00
Cory Dransfeldt
4568d69b0e
chore(robots.json): adds Panscient 2025-06-04 10:54:14 -07:00
Cory Dransfeldt
03831a7eb5
chore(robots.json): adds Quillbot 2025-06-04 10:46:58 -07:00
dark-visitors
2b5a59a303 Update from Dark Visitors 2025-06-04 01:00:07 +00:00
ai.robots.txt
3efabc603d Merge pull request #141 from Ivan-Chupin/patch-1
Add SBIntuitionsBot
2025-06-03 23:28:48 +00:00
Cory Dransfeldt
b35f9a31d7
Merge pull request #141 from Ivan-Chupin/patch-1
Add SBIntuitionsBot
2025-06-03 16:28:36 -07:00
Ivan Chupin
8f75f4a2f5
Add SBIntuitionsBot 2025-06-04 03:48:42 +05:00
ai.robots.txt
080946c360 Merge pull request #140 from ai-robots-txt/yandex-bots
chore(robots.json): adds YandexAdditional crawlers
2025-06-03 19:51:25 +00:00
Glyn Normington
7eec033cad
Merge pull request #140 from ai-robots-txt/yandex-bots
chore(robots.json): adds YandexAdditional crawlers
2025-06-03 20:51:14 +01:00
Cory Dransfeldt
3187fd8a32
chore(robots.json): adds YandexAdditional crawlers 2025-06-03 12:41:57 -07:00
ai.robots.txt
d239e7e5ad Merge pull request #139 from ai-robots-txt/workflow-fix
chore(ai_robots_update.yml): correct workflow by revising git flags + adding guard
2025-06-03 01:52:35 +00:00
Glyn Normington
9dbf34010a
Merge pull request #139 from ai-robots-txt/workflow-fix
chore(ai_robots_update.yml): correct workflow by revising git flags + adding guard
2025-06-03 02:52:23 +01:00
dark-visitors
87016d1504 Update from Dark Visitors 2025-06-03 01:00:29 +00:00
Cory Dransfeldt
899ce01c55
chore(ai_robots_update.yml): correct workflow by revising git flags + adding guard 2025-06-02 14:56:09 -07:00
Glyn Normington
4af776f0a0
Merge pull request #136 from ai-robots-txt/imgproxy-revert
chore(robots.json): revert "adds imgproxy crawler"
2025-06-02 20:21:10 +01:00
Cory Dransfeldt
1dd66b6969
Revert "chore(robots.json): adds imgproxy crawler"
This reverts commit b65f45e408.
2025-06-02 11:53:06 -07:00
Cory Dransfeldt
814df6b9a0
Merge pull request #134 from not-not-the-imp/patch-1
Add AndiBot and PhindBot
2025-05-31 16:03:16 -07:00
Cory Dransfeldt
268922f8f2
Update robots.json 2025-05-31 16:02:05 -07:00
Cory Dransfeldt
4259b25ccc
Update robots.json 2025-05-31 16:01:09 -07:00
Cory Dransfeldt
d22b9ec51a
Update robots.json 2025-05-31 16:00:13 -07:00
imp
3e8edd083e
Add AndiBot and PhindBot
Fixes #75
2025-05-23 13:03:49 +01:00
ai.robots.txt
093ab81d78 Update from Dark Visitors 2025-05-23 00:58:57 +00:00
dark-visitors
7bf7f9164d Update from Dark Visitors 2025-05-22 00:58:45 +00:00
ai.robots.txt
fedb658cc0 Merge pull request #133 from ai-robots-txt/wpbot
chore(robots.json): adds wpbot
2025-05-21 21:06:05 +00:00
Glyn Normington
851eabe059
Merge pull request #133 from ai-robots-txt/wpbot
chore(robots.json): adds wpbot
2025-05-21 22:05:51 +01:00
ai.robots.txt
7c5389f4a0 Merge pull request #98 from kylebuckingham/main
Updating Claude Bots
2025-05-21 19:00:23 +00:00
Cory Dransfeldt
af597586b6
Merge pull request #98 from kylebuckingham/main
Updating Claude Bots
2025-05-21 12:00:11 -07:00
Cory Dransfeldt
b1d9a60a38
chore(robots.json): adds wpbot 2025-05-21 11:40:33 -07:00
ai.robots.txt
1c2acd75b7 Merge pull request #126 from ai-robots-txt/mistral-bot
chore(robots.json): adds MistralAI-User/1.0 crawler
2025-05-21 15:27:26 +00:00
Glyn Normington
202d3c3b9a
Merge pull request #126 from ai-robots-txt/mistral-bot
chore(robots.json): adds MistralAI-User/1.0 crawler
2025-05-21 16:27:14 +01:00
Glyn Normington
0a78fe1e76
Merge pull request #132 from ai-robots-txt/crawler-policy-update
chore(README): updates the opening line of our README to clarify the types of agents we block
2025-05-21 15:13:35 +01:00
Cory Dransfeldt
8b151b2cdc
Update README.md
Co-authored-by: Glyn Normington <glyn.normington@gmail.com>
2025-05-21 06:52:36 -07:00
Cory Dransfeldt
8a8001cbec
chore(README): updates the opening line of our README to clarify the types of agents we block 2025-05-20 13:55:25 -07:00
Glyn Normington
fe1267e290
Merge pull request #131 from Mihitoko/mention-x-robots-tag-for-bing
Mention X-Robots-Tag header as alternative for bing
2025-05-20 07:52:32 +01:00
Mihitoko
9297c7dfa3
Mention X-Robots-Tag header as alternative for bing 2025-05-20 00:10:05 +02:00
dark-visitors
7a2e6cba52 Update from Dark Visitors 2025-05-17 00:57:28 +00:00
Falko Zurell
6c3ae6eb20
Merge branch 'ai-robots-txt:main' into main 2025-05-16 13:52:02 +02:00
Falko Zurell
684d11d889 moved traefik manual setup into docs
moved the traefik manual setup into the docs directory and linked to it from the README.md
2025-05-16 13:50:58 +02:00
ai.robots.txt
dd1ed174b7 Merge pull request #129 from ai-robots-txt/google-cloudvertexbot
chore(robots.json): adds Google-CloudVertexBot
2025-05-16 11:35:15 +00:00
Glyn Normington
89c0fbaf86
Merge pull request #129 from ai-robots-txt/google-cloudvertexbot
chore(robots.json): adds Google-CloudVertexBot
2025-05-16 12:35:04 +01:00
Falko Zurell
9b5f75e2f3
Create traefik-manual-setup.md 2025-05-16 13:27:31 +02:00
Cory Dransfeldt
ca918a963f
chore(robots.json): adds Google-CloudVertexBot 2025-05-15 21:16:49 -07:00
Cory Dransfeldt
5fba0b746d
chore(robots.json): adds MistralAI-User/1.0 crawler 2025-05-15 20:45:20 -07:00
dark-visitors
16d1de7094 Update from Dark Visitors 2025-05-16 00:59:08 +00:00
Glyn Normington
73f6f67adf
Merge pull request #125 from holysoles/lint_robots_json
lint robots.json during pull requests
2025-05-15 17:26:15 +01:00
Patrick Evans
498aa50760 lint robots.json during pull requests 2025-05-15 11:15:25 -05:00
ai.robots.txt
1c470babbe Merge pull request #123 from joehoyle/patch-1
Fix JSON syntax error
2025-05-15 16:12:30 +00:00
Adam Newbold
84d63916d2
Merge pull request #123 from joehoyle/patch-1
Fix JSON syntax error
2025-05-15 12:12:21 -04:00
Joe Hoyle
0c56b96fd9
Fix JSON syntax error 2025-05-15 11:26:47 -04:00
Cory Dransfeldt
28e69e631b
Merge pull request #122 from ai-robots-txt/qualified-bot
chore(robots.json): adds QualifiedBot crawler
2025-05-15 07:17:51 -07:00
Cory Dransfeldt
9539256cb3
chore(robots.json): adds QualifiedBot crawler 2025-05-15 07:16:07 -07:00
Cory Dransfeldt
9659c88b0c
Merge pull request #121 from solution-libre/add-traefik-plugin
Add Traefik plugin to the README.md file
2025-05-14 16:45:34 -07:00
Florent Poinsaut
c66d180295
Merge branch 'main' into add-traefik-plugin 2025-05-14 22:06:56 +02:00
Glyn Normington
9a9b1b41c0
Merge pull request #119 from ai-robots-txt/bing-ai-opt-out-instructions
Bing AI opt-out instructions
2025-05-14 19:18:20 +01:00
Florent Poinsaut
b4610a725c Add Traefik plugin 2025-05-14 14:11:56 +02:00
Cory Dransfeldt
36a52a88d8
Bing AI opt-out instructions 2025-05-12 20:20:18 -07:00
ai.robots.txt
678380727e Merge pull request #115 from glyn/syntax
Fix Python syntax error
2025-05-01 10:29:06 +00:00
Glyn Normington
fb8188c49d
Merge pull request #115 from glyn/syntax
Fix Python syntax error
2025-05-01 11:28:54 +01:00
Glyn Normington
ec995cd686 Fix Python syntax error 2025-05-01 11:27:40 +01:00
Crazyroostereye
1310dbae46
Added a Caddyfile converter (#110)
Co-authored-by: Julian Beittel <julian@beittel.net>
Co-authored-by: Glyn Normington <work@underlap.org>
2025-05-01 11:21:32 +01:00
Glyn Normington
91a88e2fa8
Merge pull request #113 from rwijnen-um/feature/haproxy
HAProxy converter added.
2025-04-28 09:00:16 +01:00
Rik Wijnen
a4a9f2ac2b Tests for HAProxy file added. 2025-04-28 09:30:26 +02:00
Rik Wijnen
66da70905f Fixed incorrect English sentence. 2025-04-28 09:09:40 +02:00
Rik Wijnen
50e739dd73 HAProxy converter added. 2025-04-28 08:51:02 +02:00
ai.robots.txt
c6c7f1748f Update from Dark Visitors 2025-04-26 00:55:12 +00:00
dark-visitors
934ac7b318 Update from Dark Visitors 2025-04-25 00:56:57 +00:00
ai.robots.txt
4654e14e9c Merge pull request #112 from maiavixen/main
Fixed meta-external* being titlecase, and removed period for consistency
2025-04-24 07:00:34 +00:00
Glyn Normington
9bf31fbca8
Merge pull request #112 from maiavixen/main
Fixed meta-external* being titlecase, and removed period for consistency
2025-04-24 08:00:24 +01:00
maia
9d846ced45
Update robots.json
Lowercase meta-external* as that was not technically the UA for the bots, also removed a period in the "respect" for consistency
2025-04-24 04:08:20 +02:00
dark-visitors
8d25a424d9 Update from Dark Visitors 2025-04-23 00:56:52 +00:00
ai.robots.txt
bbec639c14 Merge pull request #109 from dennislee1/patch-1
AI bots to consider adding
2025-04-22 14:50:26 +00:00
Cory Dransfeldt
422cf9e29b
Merge pull request #109 from dennislee1/patch-1
AI bots to consider adding
2025-04-22 07:50:14 -07:00
Dennis Lee
33c5ce1326
Update robots.json
Updated robots list with five new proposed AI bots:

aiHitBot
Cotoyogi
Factset_spyderbot
FirecrawlAgent
TikTokSpider
2025-04-21 18:55:11 +01:00
Cory Dransfeldt
774b1ddf52
Merge pull request #107 from glyn/sponsorship
Clarify our position on sponsorship
2025-04-18 11:40:06 -07:00
Glyn Normington
b1856e6988 Donations 2025-04-18 18:40:44 +01:00
Glyn Normington
d05ede8fe1 Clarify our position on sponsorship
Some firms, including those with .ai domains, have
offered to sponsor this project. So make our position
clear.
2025-04-18 17:46:56 +01:00
Kyle Buckingham
fd41de8522
Update robots.json
Co-authored-by: Glyn Normington <work@underlap.org>
2025-04-16 16:43:03 -07:00
Kyle Buckingham
4a6f37d727
Update robots.json
Co-authored-by: Glyn Normington <work@underlap.org>
2025-04-16 16:42:58 -07:00
ai.robots.txt
e0cdb278fb Update from Dark Visitors 2025-04-16 00:57:11 +00:00
dark-visitors
a96e330989 Update from Dark Visitors 2025-04-15 00:57:01 +00:00
Cory Dransfeldt
156e6baa09
Merge pull request #105 from jsheard/patch-1
Include "AI Agents" from Dark Visitors
2025-04-14 10:08:38 -07:00
Joshua Sheard
d9f882a9b2
Include "AI Agents" from Dark Visitors 2025-04-14 15:46:01 +01:00
dark-visitors
305188b2e7 Update from Dark Visitors 2025-04-11 00:55:52 +00:00
ai.robots.txt
4a764bba18 Merge pull request #102 from ai-robots-txt/imgproxy-bot
chore(robots.json): adds imgproxy crawler
2025-04-10 19:22:34 +00:00
Cory Dransfeldt
a891ad7213
Merge pull request #102 from ai-robots-txt/imgproxy-bot
chore(robots.json): adds imgproxy crawler
2025-04-10 12:22:23 -07:00
Cory Dransfeldt
b65f45e408
chore(robots.json): adds imgproxy crawler 2025-04-10 10:12:51 -07:00
Glyn Normington
49e58b1573
Merge pull request #100 from fbartho/fb/fix-perplexity-users
Fix html-mangled hyphen in 'Perplexity-Users' bot name
2025-04-05 17:32:19 +01:00
Frederic Barthelemy
c6f308cbd0
PR Feedback: log special-case, comment consistency 2025-04-05 09:01:52 -07:00
Frederic Barthelemy
5f5a89c38c
Fix html-mangled hyphen in Perplexity-Users
Fixes: #99
2025-04-04 17:34:14 -07:00
Frederic Barthelemy
6b0349f37d
fix python complaining about f-string syntax
```
python code/tests.py
Traceback (most recent call last):
  File "/Users/fbarthelemy/Code/ai.robots.txt/code/tests.py", line 7, in <module>
    from robots import json_to_txt, json_to_table, json_to_htaccess, json_to_nginx
  File "/Users/fbarthelemy/Code/ai.robots.txt/code/robots.py", line 144
    return f"({"|".join(map(re.escape, lst))})"
                ^
SyntaxError: f-string: expecting '}'
```
2025-04-04 15:20:30 -07:00
Kyle Buckingham
8dc36aa2e2
Update robots.txt 2025-04-01 15:23:28 -07:00
Kyle Buckingham
ae8f74c10c
Update robots.json 2025-04-01 15:22:04 -07:00
ai.robots.txt
5b8650b99b Update from Dark Visitors 2025-03-29 00:54:10 +00:00
dark-visitors
c249de99a3 Update from Dark Visitors 2025-03-28 00:54:28 +00:00
Cory Dransfeldt
ec18af7624
Revert "Merge pull request #91 from deyigifts/perplexity-user"
This reverts commit 68d1d93714.
2025-03-27 12:51:22 -07:00
ai.robots.txt
6851413c52 Merge pull request #94 from ThomasLeister/feature/implement-nginx-configuration-snippet-export
Implement Nginx configuration snippet export
2025-03-27 19:49:15 +00:00
Glyn Normington
dba03d809c
Merge pull request #94 from ThomasLeister/feature/implement-nginx-configuration-snippet-export
Implement Nginx configuration snippet export
2025-03-27 19:49:05 +00:00
ai.robots.txt
68d1d93714 Merge pull request #91 from deyigifts/perplexity-user
Update perplexity bots
2025-03-27 19:29:30 +00:00
Cory Dransfeldt
1183187be9
Merge pull request #91 from deyigifts/perplexity-user
Update perplexity bots
2025-03-27 12:29:21 -07:00
Thomas Leister
7c3b5a2cb2
Add tests for Nginx config generator 2025-03-27 18:28:21 +01:00
Thomas Leister
4f3f4cd0dd
Add assembled version of nginx-block-ai-bots.conf file 2025-03-27 12:43:36 +01:00
Thomas Leister
5a312c5f4d
Mention Nginx config feature in README 2025-03-27 12:43:29 +01:00
Thomas Leister
da85207314
Implement new function "json_to_nginx" which outputs an Nginx
configuration snippet
2025-03-27 12:27:09 +01:00
deyigifts
6ecfcdfcbf
Update perplexity bot
Update based on perplexity bot docs
2025-03-24 14:16:57 +08:00
Cory Dransfeldt
5e7c3c432f
Merge pull request #83 from glyn/81-doc-testing
Document testing in README
2025-02-19 09:19:44 -08:00
Glyn Normington
9f41d4c11c
Merge pull request #84 from sideeffect42/tests-workflow
Add run-tests workflow
2025-02-18 19:42:55 +00:00
Dennis Camera
8a74896333 Add workflow to run tests on pull request or push to main 2025-02-18 20:30:27 +01:00
Glyn Normington
1d55a205e4 Document testing in README
Fixes: https://github.com/ai-robots-txt/ai.robots.txt/issues/81
2025-02-18 16:49:08 +00:00
Glyn Normington
8494a7fcaa
Merge pull request #80 from sideeffect42/htaccess-allow-robots_txt
.htaccess: Allow robots access to `/robots.txt`
2025-02-18 16:42:36 +00:00
Dennis Camera
c7c1e7b96f robots.py: Make executable 2025-02-18 12:55:17 +01:00
Dennis Camera
17b826a6d3 Update tests and convert to stock unittest
For these simple tests Python's built-in unittest framework is more than enough.
No additional dependencies are required.

Added some more test cases with "special" characters to test the escaping code
better.
2025-02-18 12:55:15 +01:00
Dennis Camera
0bd3fa63b8 table-of-bot-metrics.md: Escape robot names for Markdown table
Some characters which could occur in a crawler's name have a special meaning in
Markdown. They are escaped to prevent them from having unintended side effects.

The escaping is only applied to the first (Name) column of the table. The rest
of the columns is expected to already be Markdown encoded in robots.json.
2025-02-18 12:53:27 +01:00
Dennis Camera
a884a2afb9 .htaccess: Make regex in RewriteCond safe
Improve the regular expression by removing unneeded anchors and
escaping special characters (not just space) to prevent false positives
or a misbehaving rewrite rule.
2025-02-18 12:53:22 +01:00
Dennis Camera
c0d418cd87 .htaccess: Allow robots access to /robots.txt 2025-02-18 12:49:29 +01:00
dark-visitors
abfd6dfcd1 Update from Dark Visitors 2025-02-17 00:53:32 +00:00
ai.robots.txt
693289bb29 chore: add Brightbot 1.0 2025-02-16 21:37:52 +00:00
Cory Dransfeldt
a9ec4ffa6f
chore: add Brightbot 1.0 2025-02-16 13:36:39 -08:00
Glyn Normington
03aa829913
Merge pull request #79 from always-be-testing/main
List of AI bots Cloudflare considers "Verified"
2025-02-16 04:33:40 +00:00
always-be-testing
5b13c2e504
add more concise message about verified bots
Co-authored-by: Glyn Normington <work@underlap.org>
2025-02-15 11:22:10 -05:00
always-be-testing
af87b85d7f include return after heading 2025-02-14 12:39:08 -05:00
always-be-testing
f99339922f grammar update and include syntax for verified bot condition 2025-02-14 12:36:33 -05:00
always-be-testing
e396a2ec78 forgot to include heading 2025-02-14 12:31:20 -05:00
always-be-testing
261a2b83b9 update README to inclide list of ai bots Cloudflare considers verified 2025-02-14 12:26:19 -05:00
dark-visitors
bebffccc0c Update from Dark Visitors 2025-02-02 00:52:50 +00:00
ai.robots.txt
89d4c6e5ca Merge pull request #73 from nisbet-hubbard/patch-8
Actually block Semrush’s AI tools
2025-02-01 10:51:01 +00:00
Glyn Normington
f9e2c5810b
Merge pull request #73 from nisbet-hubbard/patch-8
Actually block Semrush’s AI tools
2025-02-01 10:50:50 +00:00
nisbet-hubbard
05b79b8a58
Update robots.json 2025-01-27 19:41:03 +08:00
dark-visitors
9c060dee1c Update from Dark Visitors 2025-01-21 00:49:22 +00:00
ai.robots.txt
6c552a3daa Merge pull request #71 from jsheard/patch-1
Add Crawlspace
2025-01-20 17:45:42 +00:00
Glyn Normington
f621fb4852
Merge pull request #71 from jsheard/patch-1
Add Crawlspace
2025-01-20 17:45:29 +00:00
Joshua Sheard
7427d96bac
Update robots.json
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-20 10:59:02 +00:00
Glyn Normington
81cc81b35e
Merge pull request #68 from MassiminoilTrace/main
Implementing automatic htaccess generation
2025-01-20 07:33:54 +00:00
Massimo Gismondi
4f03818280 Removed if condition and added a little comments 2025-01-20 06:51:06 +01:00
Massimo Gismondi
a9956f7825 Removed additional sections 2025-01-20 06:50:48 +01:00
Massimo Gismondi
33c38ee70b
Update README.md
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-20 06:28:32 +01:00
Massimo Gismondi
52241bdca6
Update README.md
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-20 06:27:56 +01:00
Massimo Gismondi
013b7abfa1
Update README.md
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-20 06:27:02 +01:00
Massimo Gismondi
70fd6c0fb1
Add mention of htaccess in readme
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-20 06:25:07 +01:00
Joshua Sheard
5aa08bc002
Add Crawlspace 2025-01-19 22:03:50 +00:00
Massimo Gismondi
d65128d10a
Removed paragraph in favour of future FAQ.md
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-18 12:41:09 +01:00
Massimo Gismondi
1cc4b59dfc
Shortened htaccess instructions
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-18 12:40:03 +01:00
Massimo Gismondi
8aee2f24bb
Fixed space in comment
Co-authored-by: Glyn Normington <work@underlap.org>
2025-01-18 12:39:07 +01:00
Massimo Gismondi
b455af66e7 Adding clarification about performance and code comment 2025-01-17 21:42:08 +01:00
Massimo Gismondi
189e75bbfd Adding usage instructions 2025-01-17 21:25:23 +01:00
Massimo Gismondi
933aa6159d Implementing htaccess generation 2025-01-07 11:02:29 +01:00
Glyn Normington
b7f908e305
Merge pull request #66 from fabianegli/patch-1
Allow Action to succeed even if no changes were made
2025-01-07 03:54:40 +00:00
ai.robots.txt
ec454b71d3 Merge pull request #67 from Nightfirecat/semrushbot
Block SemrushBot
2025-01-06 20:51:56 +00:00
Cory Dransfeldt
565dca3dc0
Merge pull request #67 from Nightfirecat/semrushbot
Block SemrushBot
2025-01-06 12:51:43 -08:00
Jordan Atwood
143f8f2285
Block SemrushBot 2025-01-06 12:34:38 -08:00
Cory Dransfeldt
8e98cc6049
Merge pull request #61 from glyn/improve-naming
Rename Python code
2025-01-06 08:10:47 -08:00
Fabian Egli
30ee957011
bail when NO changes are staged 2025-01-06 12:05:42 +01:00
Fabian Egli
83cd546470
allow Action to succeed even if no changes were made
Before, the Action would fail in case there were no changes made to any files by the converter.
2025-01-06 11:39:41 +01:00
ai.robots.txt
ca8620e28b Merge pull request #63 from glyn/push-paths
Convert robots.json more frequently
2025-01-05 05:05:20 +00:00
Glyn Normington
b9df958b39
Merge pull request #63 from glyn/push-paths
Convert robots.json more frequently
2025-01-05 05:05:01 +00:00
Glyn Normington
c01a684036 Convert robots.json more frequently
Specifically, when github workflows or code
is changed as either of these can affect the
conversion results.

Ref: https://github.com/ai-robots-txt/ai.robots.txt/issues/60
2025-01-05 05:03:50 +00:00
Glyn Normington
d2be15447c
Merge pull request #62 from ai-robots-txt/missing-dependency
Ensure dependency installed
2025-01-05 01:46:27 +00:00
Glyn Normington
996b9c678c Improve job name
The purpose of the job is to convert the JSON file
to the other files.
2025-01-04 05:28:41 +00:00
Glyn Normington
e4c12ee2f8 Rename in test code 2025-01-04 05:03:48 +00:00
Glyn Normington
3a43714908 Rename Python code
The name dark_visitors.py gives the impression that the code is entirely
related to the dark visitors website, whereas the update command relates
to dark visitors and the convert command is unrelated to dark visitors.
2025-01-04 04:55:34 +00:00
28 changed files with 2511 additions and 325 deletions

9
.editorconfig Normal file
View file

@ -0,0 +1,9 @@
root = true
[*]
end_of_line = lf
insert_final_newline = true
trim_trailing_whitespace = true
[{Caddyfile,haproxy-block-ai-bots.txt,nginx-block-ai-bots.conf}]
insert_final_newline = false

View file

@ -2,29 +2,32 @@ name: Updates for AI robots files
on: on:
schedule: schedule:
- cron: "0 0 * * *" - cron: "0 0 * * *"
workflow_dispatch:
jobs: jobs:
dark-visitors: known-agents:
runs-on: ubuntu-latest runs-on: ubuntu-latest
name: dark-visitors name: known-agents
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with: with:
fetch-depth: 2 fetch-depth: 0
- run: | - run: |
pip install beautifulsoup4 requests pip install beautifulsoup4 requests
git config --global user.name "dark-visitors" git config --global user.name "Known Agents"
git config --global user.email "dark-visitors@users.noreply.github.com" git config --global user.email "knownagents@users.noreply.github.com"
echo "Updating robots.json with data from darkvisitor.com ..." echo "Updating robots.json with data from knownagents.com ..."
python code/dark_visitors.py --update python code/robots.py --update
echo "Updating generated files from robots.json ..."
python code/robots.py --convert
echo "... done." echo "... done."
git --no-pager diff git --no-pager diff
git add -A git add -A
git diff --quiet && git diff --staged --quiet || (git commit -m "Update from Dark Visitors" && git push) if ! git diff --cached --quiet; then
git commit -m "Update from knownagents.com"
git rebase origin/main
git push
else
echo "No changes to commit."
fi
shell: bash shell: bash
call-main:
needs: dark-visitors
uses: ./.github/workflows/main.yml
secrets: inherit
with:
message: "Update from Dark Visitors"

View file

@ -8,6 +8,8 @@ on:
push: push:
paths: paths:
- 'robots.json' - 'robots.json'
- '.github/workflows/**'
- 'code/**'
branches: branches:
- "main" - "main"
@ -19,21 +21,31 @@ jobs:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with: with:
fetch-depth: 2 fetch-depth: 2
- run: | - env:
INPUT_MESSAGE: ${{ inputs.message }}
HEAD_COMMIT_MESSAGE: ${{ github.event.head_commit.message }}
run: |
pip install beautifulsoup4 pip install beautifulsoup4
git config --global user.name "ai.robots.txt" git config --global user.name "ai.robots.txt"
git config --global user.email "ai.robots.txt@users.noreply.github.com" git config --global user.email "ai.robots.txt@users.noreply.github.com"
git log -1 git log -1
git status git status
echo "Updating robots.txt and table-of-bot-metrics.md if necessary ..." echo "Updating robots.txt and table-of-bot-metrics.md if necessary ..."
python code/dark_visitors.py --convert python code/robots.py --convert
echo "... done." echo "... done."
git --no-pager diff git --no-pager diff
git add -A git add -A
if [ -n "${{ inputs.message }}" ]; then if [ -z "$(git diff --staged)" ]; then
git commit -m "${{ inputs.message }}" # To have the action run successfully, if no changes are staged, we
# manually skip the later commits because they fail with exit code 1
# and this would then display as a failure for the Action.
echo "No staged changes to commit. Skipping commit and push."
exit 0
fi
if [ -n "$INPUT_MESSAGE" ]; then
git commit -m "$INPUT_MESSAGE"
else else
git commit -m "${{ github.event.head_commit.message }}" git commit -m "$HEAD_COMMIT_MESSAGE"
fi fi
git push git push
shell: bash shell: bash

28
.github/workflows/run-tests.yml vendored Normal file
View file

@ -0,0 +1,28 @@
on:
pull_request:
branches:
- main
push:
branches:
- main
jobs:
run-tests:
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v4
with:
fetch-depth: 2
- name: Install dependencies
run: |
pip install -U requests beautifulsoup4
- name: Run tests
run: |
code/tests.py
lint-json:
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v4
- name: JQ Json Lint
run: jq . robots.json

3
.htaccess Normal file
View file

@ -0,0 +1,3 @@
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|KeenableBot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b) [NC]
RewriteRule !^/?robots\.txt$ - [F]

3
Caddyfile Normal file
View file

@ -0,0 +1,3 @@
@aibots {
header_regexp User-Agent "(?i)\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|KeenableBot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b"
}

44
FAQ.md
View file

@ -2,20 +2,36 @@
## Why should we block these crawlers? ## Why should we block these crawlers?
They're extractive, confer no benefit to the creators of data they're ingesting and also have wide-ranging negative externalities: particularly copyright abuse and environmental impact. They're extractive, confer no benefit to the creators of data they're ingesting
and also have wide-ranging negative externalities, from copyright abuse and
environmental impacts to the exploitation of labor and use in war.
### References
- [How Tech Giants Cut Corners to Harvest Data for A.I.](https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html?unlocked_article_code=1.ik0.Ofja.L21c1wyW-0xj&ugrp=m)
**[How Tech Giants Cut Corners to Harvest Data for A.I.](https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html?unlocked_article_code=1.ik0.Ofja.L21c1wyW-0xj&ugrp=m)**
> OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems. > OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems.
**[How AI copyright lawsuits could make the whole industry go extinct](https://www.theverge.com/24062159/ai-copyright-fair-use-lawsuits-new-york-times-openai-chatgpt-decoder-podcast)** - [How AI copyright lawsuits could make the whole industry go extinct](https://www.theverge.com/24062159/ai-copyright-fair-use-lawsuits-new-york-times-openai-chatgpt-decoder-podcast)
> The New York Times' lawsuit against OpenAI is part of a broader, industry-shaking copyright challenge that could define the future of AI. > The New York Times' lawsuit against OpenAI is part of a broader, industry-shaking copyright challenge that could define the future of AI.
**[Reconciling the contrasting narratives on the environmental impact of large language models](https://www.nature.com/articles/s41598-024-76682-6)** - [Reconciling the contrasting narratives on the environmental impact of large language models](https://www.nature.com/articles/s41598-024-76682-6)
> Studies have shown that the training of just one LLM can consume as much energy as five cars do across their lifetimes. The water footprint of AI is also substantial; for example, recent work has highlighted that water consumption associated with AI models involves data centers using millions of gallons of water per day for cooling. Additionally, the energy consumption and carbon emissions of AI are projected to grow quickly in the coming years [...]. > Studies have shown that the training of just one LLM can consume as much energy as five cars do across their lifetimes. The water footprint of AI is also substantial; for example, recent work has highlighted that water consumption associated with AI models involves data centers using millions of gallons of water per day for cooling. Additionally, the energy consumption and carbon emissions of AI are projected to grow quickly in the coming years [...].
**[Scientists Predict AI to Generate Millions of Tons of E-Waste](https://www.sciencealert.com/scientists-predict-ai-to-generate-millions-of-tons-of-e-waste)** - [Scientists Predict AI to Generate Millions of Tons of E-Waste](https://www.sciencealert.com/scientists-predict-ai-to-generate-millions-of-tons-of-e-waste)
> we could end up with between 1.2 million and 5 million metric tons of additional electronic waste by the end of this decade [the 2020's]. > we could end up with between 1.2 million and 5 million metric tons of additional electronic waste by the end of this decade [the 2020's].
- [Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic](https://time.com/6247678/openai-chatgpt-kenya-workers/)
> AI often relies on hidden human labor in the Global South that can often be damaging and exploitative. These invisible workers remain on the margins even as their work contributes to billion-dollar industries.
- [US Military Using Claude to Select Targets in Iran Strikes](https://futurism.com/artificial-intelligence/claude-anthropic-military-iran)
> Anthropic’s large language model, Claude, is the key “AI tool” used by US Central Command in the Middle East. Its tasks include assessing intelligence, simulated war games, and even identifying military targets — in short, helping military leaders plan attacks that have already claimed hundreds of lives.
## How do we know AI companies/bots respect `robots.txt`? ## How do we know AI companies/bots respect `robots.txt`?
The short answer is that we don't. `robots.txt` is a well-established standard, but compliance is voluntary. There is no enforcement mechanism. The short answer is that we don't. `robots.txt` is a well-established standard, but compliance is voluntary. There is no enforcement mechanism.
@ -32,6 +48,16 @@ Yes, provided the crawlers identify themselves and your application/hosting supp
Some crawlers — [such as Perplexity](https://rknight.me/blog/perplexity-ai-is-lying-about-its-user-agent/) — do not identify themselves via their user agent strings and, as such, are difficult to block. Some crawlers — [such as Perplexity](https://rknight.me/blog/perplexity-ai-is-lying-about-its-user-agent/) — do not identify themselves via their user agent strings and, as such, are difficult to block.
## Can I use `robots.json` directly in my own tooling?
You're welcome to, with a caveat. `robots.json` isn't intended as a primary
deliverable of this project — the generated configuration files are. If you consume
it yourself, note that the agent names should be matched as whole words rather than
substrings.
The generated configs wrap items in `\b(...)\b` for this reason. Without
word boundaries, agent names match inside unrelated strings (see [issue 208](https://github.com/ai-robots-txt/ai.robots.txt/issues/208) for an example).
## What can we do if a bot doesn't respect `robots.txt`? ## What can we do if a bot doesn't respect `robots.txt`?
That depends on your stack. That depends on your stack.
@ -55,3 +81,11 @@ That depends on your stack.
## How can I contribute? ## How can I contribute?
Open a pull request. It will be reviewed and acted upon appropriately. **We really appreciate contributions** — this is a community effort. Open a pull request. It will be reviewed and acted upon appropriately. **We really appreciate contributions** — this is a community effort.
## I'd like to donate money
That's kind of you, but we don't need your money. If you insist, we'd love you to make a donation to the [American Civil Liberties Union](https://www.aclu.org/), the [Disasters Emergency Committee](https://www.dec.org.uk/), or a similar organisation.
## Can my company sponsor ai.robots.txt?
No, thank you. We do not accept sponsorship of any kind. We prefer to maintain our independence. Our costs are negligible as we are entirely volunteer-based and community-driven.

104
README.md
View file

@ -1,32 +1,117 @@
# ai.robots.txt # ai.robots.txt
<img src="/assets/images/noai-logo.png" width="100" /> <img src="/assets/images/noai-logo.png" width="100" alt="No AI entry logo" />
This is an open list of web crawlers associated with AI companies and the training of LLMs to block. We encourage you to contribute to and implement this list on your own site. See [information about the listed crawlers](./table-of-bot-metrics.md) and the [FAQ](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/FAQ.md). This list contains AI-related crawlers of all types, regardless of purpose. We encourage you to contribute to and implement this list on your own site. See [information about the listed crawlers](./table-of-bot-metrics.md) and the [FAQ](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/FAQ.md).
A number of these crawlers have been sourced from [Dark Visitors](https://darkvisitors.com) and we appreciate the ongoing effort they put in to track these crawlers. A number of these crawlers have been sourced from [Known Agents](https://knownagents.com) and we appreciate the ongoing effort they put in to track these crawlers.
If you'd like to add information about a crawler to the list, please make a pull request with the bot name added to `robots.txt`, `ai.txt`, and any relevant details in `table-of-bot-metrics.md` to help people understand what's crawling. If you'd like to add an AI-related crawler to the list, please see "Contributing" below.
## Usage
This repository provides the following files:
- `robots.txt`
- `.htaccess`
- `nginx-block-ai-bots.conf`
- `Caddyfile`
- `haproxy-block-ai-bots.txt`
- `lighttpd-block-ai-bots.conf`
`robots.txt` implements the Robots Exclusion Protocol ([RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html)).
`.htaccess` may be used to configure web servers such as [Apache httpd](https://httpd.apache.org/) to return an error page when one of the listed AI crawlers sends a request to the web server.
Note that, as stated in the [httpd documentation](https://httpd.apache.org/docs/current/howto/htaccess.html), more performant methods than an `.htaccess` file exist.
`nginx-block-ai-bots.conf` implements a Nginx configuration snippet that can be included in any virtual host `server {}` block via the `include` directive.
`Caddyfile` includes a Header Regex matcher group you can copy or import into your Caddyfile, the rejection can then be handled as followed `abort @aibots`
`haproxy-block-ai-bots.txt` may be used to configure HAProxy to block AI bots. To implement it:
1. Add the file to the config directory of HAProxy
2. Add the following lines in the `frontend` section:
```
acl ai_robot hdr_sub(user-agent) -i -f /etc/haproxy/haproxy-block-ai-bots.txt
http-request deny if ai_robot
```
(Note that the path of the `haproxy-block-ai-bots.txt` may be different in your environment.)
`lighttpd-block-ai-bots.conf` can be included with `include "fragments/lighttpd-block-ai-bots.conf"` in your lighttpd configuration either globally or in any conditional section.
[Bing uses the data it crawls for AI and training, you may opt out by adding a `meta` tag to the `head` of your site.](./docs/additional-steps/bing.md)
### Related
- [Robots.txt Traefik plugin](https://plugins.traefik.io/plugins/681b2f3fba3486128fc34fae/robots-txt-plugin):
middleware plugin for [Traefik](https://traefik.io/traefik/) to automatically add rules of [robots.txt](./robots.txt)
file on-the-fly.
- Alternatively you can [manually configure Traefik](./docs/traefik-manual-setup.md) to centrally serve a static `robots.txt`.
- [Bot Ledger](https://farrelldan.github.io/ai-bot-directory/): free, static directory of verified AI crawlers with a one-click `robots.txt` and `llms.txt` generator. No signup required.
- [KI-Zugangsindex](https://peppe1337.github.io/ki-zugangsindex/): open dataset on how widely this kind of blocking is actually deployed in the German (`.de`) web, measured on a fixed panel of 600 domains so the same sites can be re-checked over time.
- [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0.
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0.
## Contributing ## Contributing
A note about contributing: updates should be added/made to `robots.json`. A GitHub action, courtesy of [Adam](https://github.com/newbold), will then generate the updated `robots.txt` and `table-of-bot-metrics.md`. Please note that AI-generated contributions are not permitted.
A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`.
You can run the tests by [installing](https://www.python.org/about/gettingstarted/) Python 3, installing the dependencies:
```console
pip install -r requirements.txt
```
and then issuing:
```console
code/tests.py
```
The `.editorconfig` file provides standard editor options for this project. See [EditorConfig](https://editorconfig.org/) for more information.
## Releasing
Admins may ship a new release `v1.n` (where `n` increments the minor version of the current release) as follows:
- Navigate to the [new release page](https://github.com/ai-robots-txt/ai.robots.txt/releases/new) on GitHub.
- Click `Select tag`, choose `Create new tag`, enter `v1.n` in the pop-up, and click `Create`.
- Enter a suitable release title (e.g. `v1.n: adds user-agent1, user-agent2`).
- Click `Generate release notes`.
- Click `Publish release`.
A GitHub action will then add the asset `robots.txt` to the release. That's it.
## Subscribe to updates ## Subscribe to updates
You can subscribe to list updates via RSS/Atom with the releases feed: You can subscribe to list updates via RSS/Atom with the releases feed:
`https://github.com/ai-robots-txt/ai.robots.txt/releases.atom`.
```
https://github.com/ai-robots-txt/ai.robots.txt/releases.atom
```
You can subscribe with [Feedly](https://feedly.com/i/subscription/feed/https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [Inoreader](https://www.inoreader.com/?add_feed=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [The Old Reader](https://theoldreader.com/feeds/subscribe?url=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [Feedbin](https://feedbin.me/?subscribe=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), or any other reader app. You can subscribe with [Feedly](https://feedly.com/i/subscription/feed/https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [Inoreader](https://www.inoreader.com/?add_feed=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [The Old Reader](https://theoldreader.com/feeds/subscribe?url=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [Feedbin](https://feedbin.me/?subscribe=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), or any other reader app.
Alternatively, you can also subscribe to new releases with your GitHub account by clicking the ⬇️ on "Watch" button at the top of this page, clicking "Custom" and selecting "Releases". Alternatively, you can also subscribe to new releases with your GitHub account by clicking the ⬇️ on "Watch" button at the top of this page, clicking "Custom" and selecting "Releases".
## License content with RSL
It is also possible to license your content to AI companies in `robots.txt` using
the [Really Simple Licensing](https://rslstandard.org) standard, with an option of
collective bargaining. A [plugin](https://github.com/Jameswlepage/rsl-wp) currently
implements RSL as well as payment processing for WordPress sites.
## Report abusive crawlers ## Report abusive crawlers
If you use [Cloudflare's hard block](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click) alongside this list, you can report abusive crawlers that don't respect `robots.txt` [here](https://docs.google.com/forms/d/e/1FAIpQLScbUZ2vlNSdcsb8LyTeSF7uLzQI96s0BKGoJ6wQ6ocUFNOKEg/viewform). If you use [Cloudflare's hard block](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click) alongside this list, you can report abusive crawlers that don't respect `robots.txt` [here](https://docs.google.com/forms/d/e/1FAIpQLScbUZ2vlNSdcsb8LyTeSF7uLzQI96s0BKGoJ6wQ6ocUFNOKEg/viewform).
But even if you don't use Cloudflare's hard block, their list of [verified bots](https://radar.cloudflare.com/traffic/verified-bots) may come in handy.
## Additional resources ## Additional resources
@ -36,3 +121,4 @@ If you use [Cloudflare's hard block](https://blog.cloudflare.com/declaring-your-
- [Blockin' bots on Netlify](https://www.jeremiak.com/blog/block-bots-netlify-edge-functions/) by Jeremia Kimelman - [Blockin' bots on Netlify](https://www.jeremiak.com/blog/block-bots-netlify-edge-functions/) by Jeremia Kimelman
- [Blocking AI web crawlers](https://underlap.org/blocking-ai-web-crawlers) by Glyn Normington - [Blocking AI web crawlers](https://underlap.org/blocking-ai-web-crawlers) by Glyn Normington
- [Block AI Bots from Crawling Websites Using Robots.txt](https://originality.ai/ai-bot-blocking) by Jonathan Gillham, Originality.AI - [Block AI Bots from Crawling Websites Using Robots.txt](https://originality.ai/ai-bot-blocking) by Jonathan Gillham, Originality.AI
- [AI Access Checker: see which AI crawlers a site's robots.txt allows or blocks](https://www.greadme.com/ai-access-checker) by Saar Twito, Greadme

View file

@ -1,183 +0,0 @@
import json
from pathlib import Path
import requests
from bs4 import BeautifulSoup
def load_robots_json():
"""Load the robots.json contents into a dictionary."""
return json.loads(Path("./robots.json").read_text(encoding="utf-8"))
def get_agent_soup():
"""Retrieve current known agents from darkvisitors.com"""
session = requests.Session()
try:
response = session.get("https://darkvisitors.com/agents")
except requests.exceptions.ConnectionError:
print(
"ERROR: Could not gather the current agents from https://darkvisitors.com/agents"
)
return
return BeautifulSoup(response.text, "html.parser")
def updated_robots_json(soup):
"""Update AI scraper information with data from darkvisitors."""
existing_content = load_robots_json()
to_include = [
"AI Assistants",
"AI Data Scrapers",
"AI Search Crawlers",
# "Archivers",
# "Developer Helpers",
# "Fetchers",
# "Intelligence Gatherers",
# "Scrapers",
# "Search Engine Crawlers",
# "SEO Crawlers",
# "Uncategorized",
"Undocumented AI Agents",
]
for section in soup.find_all("div", {"class": "agent-links-section"}):
category = section.find("h2").get_text()
if category not in to_include:
continue
for agent in section.find_all("a", href=True):
name = agent.find("div", {"class": "agent-name"}).get_text().strip()
desc = agent.find("p").get_text().strip()
default_values = {
"Unclear at this time.",
"No information provided.",
"No information.",
"No explicit frequency provided.",
}
default_value = "Unclear at this time."
# Parse the operator information from the description if possible
operator = default_value
if "operated by " in desc:
try:
operator = desc.split("operated by ", 1)[1].split(".", 1)[0].strip()
except Exception as e:
print(f"Error: {e}")
def consolidate(field: str, value: str) -> str:
# New entry
if name not in existing_content:
return value
# New field
if field not in existing_content[name]:
return value
# Unclear value
if (
existing_content[name][field] in default_values
and value not in default_values
):
return value
# Existing value
return existing_content[name][field]
existing_content[name] = {
"operator": consolidate("operator", operator),
"respect": consolidate("respect", default_value),
"function": consolidate("function", f"{category}"),
"frequency": consolidate("frequency", default_value),
"description": consolidate(
"description",
f"{desc} More info can be found at https://darkvisitors.com/agents{agent['href']}",
),
}
print(f"Total: {len(existing_content)}")
sorted_keys = sorted(existing_content, key=lambda k: k.lower())
sorted_robots = {k: existing_content[k] for k in sorted_keys}
return sorted_robots
def ingest_darkvisitors():
old_robots_json = load_robots_json()
soup = get_agent_soup()
if soup:
robots_json = updated_robots_json(soup)
print(
"robots.json is unchanged."
if robots_json == old_robots_json
else "robots.json got updates."
)
Path("./robots.json").write_text(
json.dumps(robots_json, indent=4), encoding="utf-8"
)
def json_to_txt(robots_json):
"""Compose the robots.txt from the robots.json file."""
robots_txt = "\n".join(f"User-agent: {k}" for k in robots_json.keys())
robots_txt += "\nDisallow: /\n"
return robots_txt
def json_to_table(robots_json):
"""Compose a markdown table with the information in robots.json"""
table = "| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |\n"
table += "|-----|----------|-----------------------|----------|------------------|-------------|\n"
for name, robot in robots_json.items():
table += f'| {name} | {robot["operator"]} | {robot["respect"]} | {robot["function"]} | {robot["frequency"]} | {robot["description"]} |\n'
return table
def update_file_if_changed(file_name, converter):
"""Update files if newer content is available and log the (in)actions."""
new_content = converter(load_robots_json())
old_content = Path(file_name).read_text(encoding="utf-8")
if old_content == new_content:
print(f"{file_name} is already up to date.")
else:
Path(file_name).write_text(new_content, encoding="utf-8")
print(f"{file_name} has been updated.")
def conversions():
"""Triggers the conversions from the json file."""
update_file_if_changed(file_name="./robots.txt", converter=json_to_txt)
update_file_if_changed(
file_name="./table-of-bot-metrics.md",
converter=json_to_table,
)
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser()
parser = argparse.ArgumentParser(
prog="ai-robots",
description="Collects and updates information about web scrapers of AI companies.",
epilog="One of the flags must be set.\n",
)
parser.add_argument(
"--update",
action="store_true",
help="Update the robots.json file with data from darkvisitors.com/agents",
)
parser.add_argument(
"--convert",
action="store_true",
help="Create the robots.txt and markdown table from robots.json",
)
args = parser.parse_args()
if not (args.update or args.convert):
print("ERROR: please provide one of the possible flags.")
parser.print_help()
if args.update:
ingest_darkvisitors()
if args.convert:
conversions()

330
code/robots.py Executable file
View file

@ -0,0 +1,330 @@
#!/usr/bin/env python3
import json
import re
import requests
from bs4 import BeautifulSoup
from pathlib import Path
to_include = [
"AI Agents",
"AI Assistants",
"AI Coding Agents",
"AI Data Providers",
"AI Data Scrapers",
"AI Search Crawlers",
# "Archivers",
# "Automated Agents",
# "Developer Helpers",
# "Fetchers",
# "Intelligence Gatherers",
# "Scrapers",
# "Search Engine Crawlers",
# "Security Scanners",
# "SEO Crawlers",
# "Uncategorized",
"Undocumented AI Agents",
]
default_values = {
"Unclear at this time.",
"No information provided.",
"No information.",
"No explicit frequency provided.",
}
default_value = "Unclear at this time."
def existing_key(existing_content, name: str) -> str:
"""Return the key robots.json already uses for this agent, ignoring case.
robots.txt user-agent matching is case-insensitive, so two entries whose
names differ only in case are the same crawler. knownagents.com has changed
the capitalisation of a name before now, and keying off the scraped name
added a second entry rather than updating the first, which both duplicated
the User-agent line in every generated file and left the curated operator
and respect values behind on the old key.
"""
if name in existing_content:
return name
lowered = name.lower()
for key in existing_content:
if key.lower() == lowered:
return key
return name
def consolidate(existing_content, name: str, field: str, value: str) -> str:
# New entry
if name not in existing_content:
return value
# New field
if field not in existing_content[name]:
return value
# Unclear value
if ( value not in default_values ):
return value
# Existing value
return existing_content[name][field]
def load_robots_json():
"""Load the robots.json contents into a dictionary."""
return json.loads(Path("./robots.json").read_text(encoding="utf-8"))
def get_agent_soup():
"""Retrieve current agent data from knownagents.com."""
session = requests.Session()
try:
response = session.get("https://knownagents.com/agents")
except requests.exceptions.ConnectionError:
print(
"ERROR: Could not gather the current agents from https://knownagents.com/agents"
)
return
return BeautifulSoup(response.text, "html.parser")
def updated_robots_json(soup):
"""Update AI scraper information with data from knownagents.com."""
existing_content = load_robots_json()
for section in soup.find_all("div", {"class": "agent-links-section"}):
category = section.find("h2").get_text()
if category not in to_include:
continue
for agent in section.find_all("a", href=True):
name = agent.find("div", {"class": "agent-name"}).get_text().strip()
name = clean_robot_name(name)
name = existing_key(existing_content, name)
desc_tag = agent.find("div", {"class": "description"})
if desc_tag is not None:
desc = desc_tag.get_text().strip()
else:
desc = "Description unavailable from knownagents.com"
# Parse the operator information from the description if possible
operator = default_value
if "operated by " in desc:
try:
operator = desc.split("operated by ", 1)[1].split(".", 1)[0].strip()
except Exception as e:
print(f"Error: {e}")
robot = {
"operator": consolidate(existing_content, name, "operator", operator),
"respect": consolidate(existing_content, name, "respect", default_value),
"function": consolidate(existing_content, name, "function", f"{category}"),
"frequency": consolidate(existing_content, name, "frequency", default_value),
"description": consolidate(
existing_content, name,
"description",
"{desc} More info can be found at https://knownagents.com{path}".format(
desc = desc,
path= agent['href'] if agent['href'].startswith('/') else "/".join("", "agents", agent["href"])
),
),
}
if "has_name_and_version" in existing_content.get(name, {}):
robot["has_name_and_version"] = existing_content[name][
"has_name_and_version"
]
existing_content[name] = robot
print(f"Total: {len(existing_content)}")
sorted_keys = sorted(existing_content, key=lambda k: k.lower())
sorted_robots = {k: existing_content[k] for k in sorted_keys}
return sorted_robots
def clean_robot_name(name):
""" Clean the robot name by removing some characters that were mangled by html software once. """
# This was specifically spotted in "Perplexity-User"
# Looks like a non-breaking hyphen introduced by the HTML rendering software
# Reading the source page for Perplexity: https://docs.perplexity.ai/guides/bots
# You can see the bot is listed several times as "Perplexity-User" with a normal hyphen,
# and it's only the Row-Heading that has the special hyphen
#
# Technically, there's no reason there wouldn't someday be a bot that
# actually uses a non-breaking hyphen, but that seems unlikely,
# so this solution should be fine for now.
result = re.sub(r"\u2011", "-", name)
if result != name:
print(f"\tCleaned '{name}' to '{result}' - unicode/html mangled chars normalized.")
return result
def ingest_known_agents():
old_robots_json = load_robots_json()
soup = get_agent_soup()
if soup:
robots_json = updated_robots_json(soup)
print(
"robots.json is unchanged."
if robots_json == old_robots_json
else "robots.json got updates."
)
Path("./robots.json").write_text(
json.dumps(robots_json, indent=4), encoding="utf-8"
)
def json_to_txt(robots_json):
"""Compose the robots.txt from the robots.json file."""
robots_txt = "\n".join(f"User-agent: {k}" for k in robots_json.keys())
robots_txt += "\nDisallow: /\n"
return robots_txt
def escape_md(s):
return re.sub(r"([]*\\|`(){}<>#+-.!_[])", r"\\\1", s)
def json_to_table(robots_json):
"""Compose a markdown table with the information in robots.json"""
table = "| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |\n"
table += "|------|----------|-----------------------|----------|------------------|-------------|\n"
for name, robot in robots_json.items():
table += f'| {escape_md(name)} | {robot["operator"]} | {robot["respect"]} | {robot["function"]} | {robot["frequency"]} | {robot["description"]} |\n'
return table
def list_to_pcre(robots_json):
# Python re is not 100% identical to PCRE which is used by Apache, but it
# should probably be close enough in the real world for re.escape to work.
# We additionally un-escape '-' since it only requires escaping within
# character classes (which are also escaped and prevented here) and '/'
# since this is not used as the regexp delimeter in any server software.
def escape(pattern):
pattern = re.escape(pattern)
for c in "-/":
pattern = pattern.replace(fr"\{c}", c)
return pattern
exact_agents = "|".join(map(escape, robots_json))
return f"\\b({exact_agents})\\b"
def json_to_htaccess(robot_json):
# Creates a .htaccess filter file. It uses a regular expression to filter out
# User agents that contain any of the blocked values.
# The regular expression is wrapped in parenthesis, so a leading [-!=<>] does
# not accidentally change which comparison type is used.
htaccess = "RewriteEngine On\n"
htaccess += f"RewriteCond %{{HTTP_USER_AGENT}} ({list_to_pcre(robot_json)}) [NC]\n"
htaccess += "RewriteRule !^/?robots\\.txt$ - [F]\n"
return htaccess
def json_to_nginx(robot_json):
# Creates an Nginx config file. This config snippet can be included in
# nginx server{} blocks to block AI bots.
# Exact User-agent matching is case-insensitive (RFC 9309 2.2.1), so this uses
# nginx's case-insensitive "~*" operator rather than "~".
config = f"set $block 0;\n\nif ($http_user_agent ~* {list_to_pcre(robot_json)!r}) {{\n set $block 1;\n}}\n\nif ($request_uri = '/robots.txt') {{\n set $block 0;\n}}\n\nif ($block) {{\n return 403;\n}}"
return config
def json_to_lighttpd(robot_json):
# Creates an Lighttpd config file. This config snippet can be included in
# Lighttpd configuration global or in $HTTP conditionals to block AI bots.
# single quotes (as returned by repr) are not valid string delimeters, so we
# must manually quote it end ensure no unescaped quotes are inside.
# Lighttpd's "=~" is case-sensitive by default; the inline (?i) flag makes it
# match the case-insensitive exact matching robots.txt itself requires.
escaped_quotes = ("(?i)" + list_to_pcre(robot_json)).replace('"', '\\"')
config = f'$HTTP["url"] != "/robots.txt" {{ $HTTP["user-agent"] =~ "{escaped_quotes}" {{ url.access-deny = ( "" ) }} }}'
return config
def json_to_caddy(robot_json):
# single quotes (as returned by repr) are not valid string delimeters, so we
# must manually quote it end ensure no unescaped quotes are inside.
# Same case-insensitivity note as lighttpd; Caddy's RE2 engine also honours (?i).
escaped_quotes = ("(?i)" + list_to_pcre(robot_json)).replace('"', '\\"')
caddyfile = "@aibots {\n "
caddyfile += f' header_regexp User-Agent "{escaped_quotes}"'
caddyfile += "\n}"
return caddyfile
def json_to_haproxy(robots_json):
# Creates a source file for HAProxy. Follow instructions in the README to implement it.
txt = "\n".join(robots_json.keys())
return txt
def update_file_if_changed(file_name, converter):
"""Update files if newer content is available and log the (in)actions."""
new_content = converter(load_robots_json())
filepath = Path(file_name)
# "touch" will create the file if it doesn't exist yet
filepath.touch()
old_content = filepath.read_text(encoding="utf-8")
if old_content == new_content:
print(f"{file_name} is already up to date.")
else:
Path(file_name).write_text(new_content, encoding="utf-8")
print(f"{file_name} has been updated.")
def conversions():
"""Triggers the conversions from the json file."""
update_file_if_changed(file_name="./robots.txt", converter=json_to_txt)
update_file_if_changed(
file_name="./table-of-bot-metrics.md",
converter=json_to_table,
)
update_file_if_changed(
file_name="./.htaccess",
converter=json_to_htaccess,
)
update_file_if_changed(
file_name="./nginx-block-ai-bots.conf",
converter=json_to_nginx,
)
update_file_if_changed(
file_name="./lighttpd-block-ai-bots.conf",
converter=json_to_lighttpd,
)
update_file_if_changed(
file_name="./Caddyfile",
converter=json_to_caddy,
)
update_file_if_changed(
file_name="./haproxy-block-ai-bots.txt",
converter=json_to_haproxy,
)
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser(
prog="ai-robots",
description="Collects and updates information about web scrapers of AI companies.",
epilog="One of the flags must be set.\n",
)
parser.add_argument(
"--update",
action="store_true",
help="Update the robots.json file with data from knownagents.com/agents",
)
parser.add_argument(
"--convert",
action="store_true",
help="Create the robots.txt and markdown table from robots.json",
)
args = parser.parse_args()
if not (args.update or args.convert):
print("ERROR: please provide one of the possible flags.")
parser.print_help()
if args.update:
ingest_known_agents()
if args.convert:
conversions()

View file

@ -0,0 +1,3 @@
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b) [NC]
RewriteRule !^/?robots\.txt$ - [F]

View file

@ -0,0 +1,3 @@
@aibots {
header_regexp User-Agent "(?i)\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b"
}

View file

@ -0,0 +1,47 @@
AI2Bot
Ai2Bot-Dolma
Amazonbot
anthropic-ai
Applebot
Applebot-Extended
Bytespider
CCBot
ChatGPT-User
Claude-Web
ClaudeBot
cohere-ai
Diffbot
FacebookBot
facebookexternalhit
FriendlyCrawler
Google-Extended
GoogleOther
GoogleOther-Image
GoogleOther-Video
GPTBot
iaskspider/2.0
ICC-Crawler
ImagesiftBot
img2dataset
ISSCyberRiskCrawler
Kangaroo Bot
Meta-ExternalAgent
Meta-ExternalFetcher
OAI-SearchBot
omgili
omgilibot
Perplexity-User
PerplexityBot
PetalBot
Scrapy
Sidetrade indexer bot
Timpibot
VelenPublicWebCrawler
Webzio-Extended
YouBot
crawler.with.dots
star***crawler
Is this a crawler?
a[mazing]{42}(robot)
2^32$
curl|sudo bash

View file

@ -0,0 +1 @@
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "(?i)\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b" { url.access-deny = ( "" ) } }

View file

@ -0,0 +1,13 @@
set $block 0;
if ($http_user_agent ~* '\\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\\ indexer\\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\\.with\\.dots|star\\*\\*\\*crawler|Is\\ this\\ a\\ crawler\\?|a\\[mazing\\]\\{42\\}\\(robot\\)|2\\^32\\$|curl\\|sudo\\ bash)\\b') {
set $block 1;
}
if ($request_uri = '/robots.txt') {
set $block 0;
}
if ($block) {
return 403;
}

View file

@ -32,7 +32,7 @@
"respect": "Unclear at this time.", "respect": "Unclear at this time.",
"function": "AI Search Crawlers", "function": "AI Search Crawlers",
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://darkvisitors.com/agents/agents/applebot" "description": "Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://knownagents.com/agents/applebot"
}, },
"Applebot-Extended": { "Applebot-Extended": {
"operator": "[Apple](https://support.apple.com/en-us/119829#datausage)", "operator": "[Apple](https://support.apple.com/en-us/119829#datausage)",
@ -186,7 +186,7 @@
"respect": "Unclear at this time.", "respect": "Unclear at this time.",
"function": "AI Data Scrapers", "function": "AI Data Scrapers",
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://darkvisitors.com/agents/agents/kangaroo-bot" "description": "Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot"
}, },
"Meta-ExternalAgent": { "Meta-ExternalAgent": {
"operator": "[Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)", "operator": "[Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)",
@ -200,7 +200,7 @@
"respect": "Unclear at this time.", "respect": "Unclear at this time.",
"function": "AI Assistants", "function": "AI Assistants",
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://darkvisitors.com/agents/agents/meta-externalfetcher" "description": "Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://knownagents.com/agents/meta-externalfetcher"
}, },
"OAI-SearchBot": { "OAI-SearchBot": {
"operator": "[OpenAI](https://openai.com)", "operator": "[OpenAI](https://openai.com)",
@ -223,6 +223,13 @@
"operator": "[Webz.io](https://webz.io/)", "operator": "[Webz.io](https://webz.io/)",
"respect": "[Yes](https://web.archive.org/web/20170704003301/http://omgili.com/Crawler.html)" "respect": "[Yes](https://web.archive.org/web/20170704003301/http://omgili.com/Crawler.html)"
}, },
"Perplexity-User": {
"operator": "[Perplexity](https://www.perplexity.ai/)",
"respect": "[No](https://docs.perplexity.ai/guides/bots)",
"function": "Used to answer queries at the request of users.",
"frequency": "Only when prompted by a user.",
"description": "Visit web pages to help provide an accurate answer and include links to the page in Perplexity response."
},
"PerplexityBot": { "PerplexityBot": {
"operator": "[Perplexity](https://www.perplexity.ai/)", "operator": "[Perplexity](https://www.perplexity.ai/)",
"respect": "[No](https://www.macstories.net/stories/wired-confirms-perplexity-is-bypassing-efforts-by-websites-to-block-its-web-crawler/)", "respect": "[No](https://www.macstories.net/stories/wired-confirms-perplexity-is-bypassing-efforts-by-websites-to-block-its-web-crawler/)",
@ -270,7 +277,7 @@
"respect": "Unclear at this time.", "respect": "Unclear at this time.",
"function": "AI Data Scrapers", "function": "AI Data Scrapers",
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://darkvisitors.com/agents/agents/webzio-extended" "description": "Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://knownagents.com/agents/webzio-extended"
}, },
"YouBot": { "YouBot": {
"operator": "[You](https://about.you.com/youchat/)", "operator": "[You](https://about.you.com/youchat/)",
@ -278,5 +285,47 @@
"function": "Scrapes data for search engine and LLMs.", "function": "Scrapes data for search engine and LLMs.",
"frequency": "No information.", "frequency": "No information.",
"description": "Retrieves data used for You.com web search engine and LLMs." "description": "Retrieves data used for You.com web search engine and LLMs."
},
"crawler.with.dots": {
"operator": "Test suite",
"respect": "No",
"function": "To ensure the code works correctly.",
"frequency": "No information.",
"description": "When used in the .htaccess regular expression dots need to be escaped."
},
"star***crawler": {
"operator": "Test suite",
"respect": "No",
"function": "To ensure the code works correctly.",
"frequency": "No information.",
"description": "When used in the .htaccess regular expression stars need to be escaped."
},
"Is this a crawler?": {
"operator": "Test suite",
"respect": "No",
"function": "To ensure the code works correctly.",
"frequency": "No information.",
"description": "When used in the .htaccess regular expression spaces and question marks need to be escaped."
},
"a[mazing]{42}(robot)": {
"operator": "Test suite",
"respect": "No",
"function": "To ensure the code works correctly.",
"frequency": "No information.",
"description": "When used in the .htaccess regular expression parantheses, braces, etc. need to be escaped."
},
"2^32$": {
"operator": "Test suite",
"respect": "No",
"function": "To ensure the code works correctly.",
"frequency": "No information.",
"description": "When used in the .htaccess regular expression RE anchor characters need to be escaped."
},
"curl|sudo bash": {
"operator": "Test suite",
"respect": "No",
"function": "To ensure the code works correctly.",
"frequency": "No information.",
"description": "When used in the .htaccess regular expression pipes need to be escaped."
} }
} }

View file

@ -30,6 +30,7 @@ User-agent: Meta-ExternalFetcher
User-agent: OAI-SearchBot User-agent: OAI-SearchBot
User-agent: omgili User-agent: omgili
User-agent: omgilibot User-agent: omgilibot
User-agent: Perplexity-User
User-agent: PerplexityBot User-agent: PerplexityBot
User-agent: PetalBot User-agent: PetalBot
User-agent: Scrapy User-agent: Scrapy
@ -38,4 +39,10 @@ User-agent: Timpibot
User-agent: VelenPublicWebCrawler User-agent: VelenPublicWebCrawler
User-agent: Webzio-Extended User-agent: Webzio-Extended
User-agent: YouBot User-agent: YouBot
User-agent: crawler.with.dots
User-agent: star***crawler
User-agent: Is this a crawler?
User-agent: a[mazing]{42}(robot)
User-agent: 2^32$
User-agent: curl|sudo bash
Disallow: / Disallow: /

View file

@ -1,42 +1,49 @@
| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description | | Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |
|-----|----------|-----------------------|----------|------------------|-------------| |------|----------|-----------------------|----------|------------------|-------------|
| AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. | | AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
| Ai2Bot-Dolma | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. | | Ai2Bot\-Dolma | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
| Amazonbot | Amazon | Yes | Service improvement and enabling answers for Alexa users. | No information provided. | Includes references to crawled website when surfacing answers via Alexa; does not clearly outline other uses. | | Amazonbot | Amazon | Yes | Service improvement and enabling answers for Alexa users. | No information provided. | Includes references to crawled website when surfacing answers via Alexa; does not clearly outline other uses. |
| anthropic-ai | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | anthropic\-ai | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
| Applebot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://darkvisitors.com/agents/agents/applebot | | Applebot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://knownagents.com/agents/applebot |
| Applebot-Extended | [Apple](https://support.apple.com/en-us/119829#datausage) | Yes | Powers features in Siri, Spotlight, Safari, Apple Intelligence, and others. | Unclear at this time. | Apple has a secondary user agent, Applebot-Extended ... [that is] used to train Apple's foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools. | | Applebot\-Extended | [Apple](https://support.apple.com/en-us/119829#datausage) | Yes | Powers features in Siri, Spotlight, Safari, Apple Intelligence, and others. | Unclear at this time. | Apple has a secondary user agent, Applebot-Extended ... [that is] used to train Apple's foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools. |
| Bytespider | ByteDance | No | LLM training. | Unclear at this time. | Downloads data to train LLMS, including ChatGPT competitors. | | Bytespider | ByteDance | No | LLM training. | Unclear at this time. | Downloads data to train LLMS, including ChatGPT competitors. |
| CCBot | [Common Crawl Foundation](https://commoncrawl.org) | [Yes](https://commoncrawl.org/ccbot) | Provides open crawl dataset, used for many purposes, including Machine Learning/AI. | Monthly at present. | Web archive going back to 2008. [Cited in thousands of research papers per year](https://commoncrawl.org/research-papers). | | CCBot | [Common Crawl Foundation](https://commoncrawl.org) | [Yes](https://commoncrawl.org/ccbot) | Provides open crawl dataset, used for many purposes, including Machine Learning/AI. | Monthly at present. | Web archive going back to 2008. [Cited in thousands of research papers per year](https://commoncrawl.org/research-papers). |
| ChatGPT-User | [OpenAI](https://openai.com) | Yes | Takes action based on user prompts. | Only when prompted by a user. | Used by plugins in ChatGPT to answer queries based on user input. | | ChatGPT\-User | [OpenAI](https://openai.com) | Yes | Takes action based on user prompts. | Only when prompted by a user. | Used by plugins in ChatGPT to answer queries based on user input. |
| Claude-Web | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | Claude\-Web | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
| ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
| cohere-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. | | cohere\-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. |
| Diffbot | [Diffbot](https://www.diffbot.com/) | At the discretion of Diffbot users. | Aggregates structured web data for monitoring and AI model training. | Unclear at this time. | Diffbot is an application used to parse web pages into structured data; this data is used for monitoring or AI model training. | | Diffbot | [Diffbot](https://www.diffbot.com/) | At the discretion of Diffbot users. | Aggregates structured web data for monitoring and AI model training. | Unclear at this time. | Diffbot is an application used to parse web pages into structured data; this data is used for monitoring or AI model training. |
| FacebookBot | Meta/Facebook | [Yes](https://developers.facebook.com/docs/sharing/bot/) | Training language models | Up to 1 page per second | Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. | | FacebookBot | Meta/Facebook | [Yes](https://developers.facebook.com/docs/sharing/bot/) | Training language models | Up to 1 page per second | Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. |
| facebookexternalhit | Meta/Facebook | [Yes](https://developers.facebook.com/docs/sharing/bot/) | No information. | Unclear at this time. | Unclear at this time. | | facebookexternalhit | Meta/Facebook | [Yes](https://developers.facebook.com/docs/sharing/bot/) | No information. | Unclear at this time. | Unclear at this time. |
| FriendlyCrawler | Unknown | [Yes](https://imho.alex-kunz.com/2024/01/25/an-update-on-friendly-crawler) | We are using the data from the crawler to build datasets for machine learning experiments. | Unclear at this time. | Unclear who the operator is; but data is used for training/machine learning. | | FriendlyCrawler | Unknown | [Yes](https://imho.alex-kunz.com/2024/01/25/an-update-on-friendly-crawler) | We are using the data from the crawler to build datasets for machine learning experiments. | Unclear at this time. | Unclear who the operator is; but data is used for training/machine learning. |
| Google-Extended | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | LLM training. | No information. | Used to train Gemini and Vertex AI generative APIs. Does not impact a site's inclusion or ranking in Google Search. | | Google\-Extended | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | LLM training. | No information. | Used to train Gemini and Vertex AI generative APIs. Does not impact a site's inclusion or ranking in Google Search. |
| GoogleOther | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." | | GoogleOther | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." |
| GoogleOther-Image | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." | | GoogleOther\-Image | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." |
| GoogleOther-Video | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." | | GoogleOther\-Video | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." |
| GPTBot | [OpenAI](https://openai.com) | Yes | Scrapes data to train OpenAI's products. | No information. | Data is used to train current and future models, removed paywalled data, PII and data that violates the company's policies. | | GPTBot | [OpenAI](https://openai.com) | Yes | Scrapes data to train OpenAI's products. | No information. | Data is used to train current and future models, removed paywalled data, PII and data that violates the company's policies. |
| iaskspider/2.0 | iAsk | No | Crawls sites to provide answers to user queries. | Unclear at this time. | Used to provide answers to user queries. | | iaskspider/2\.0 | iAsk | No | Crawls sites to provide answers to user queries. | Unclear at this time. | Used to provide answers to user queries. |
| ICC-Crawler | [NICT](https://nict.go.jp) | Yes | Scrapes data to train and support AI technologies. | No information. | Use the collected data for artificial intelligence technologies; provide data to third parties, including commercial companies; those companies can use the data for their own business. | | ICC\-Crawler | [NICT](https://nict.go.jp) | Yes | Scrapes data to train and support AI technologies. | No information. | Use the collected data for artificial intelligence technologies; provide data to third parties, including commercial companies; those companies can use the data for their own business. |
| ImagesiftBot | [ImageSift](https://imagesift.com) | [Yes](https://imagesift.com/about) | ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support our suite of web intelligence products | No information. | Once images and text are downloaded from a webpage, ImageSift analyzes this data from the page and stores the information in an index. Our web intelligence products use this index to enable search and retrieval of similar images. | | ImagesiftBot | [ImageSift](https://imagesift.com) | [Yes](https://imagesift.com/about) | ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support our suite of web intelligence products | No information. | Once images and text are downloaded from a webpage, ImageSift analyzes this data from the page and stores the information in an index. Our web intelligence products use this index to enable search and retrieval of similar images. |
| img2dataset | [img2dataset](https://github.com/rom1504/img2dataset) | Unclear at this time. | Scrapes images for use in LLMs. | At the discretion of img2dataset users. | Downloads large sets of images into datasets for LLM training or other purposes. | | img2dataset | [img2dataset](https://github.com/rom1504/img2dataset) | Unclear at this time. | Scrapes images for use in LLMs. | At the discretion of img2dataset users. | Downloads large sets of images into datasets for LLM training or other purposes. |
| ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. | | ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. |
| Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://darkvisitors.com/agents/agents/kangaroo-bot | | Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot |
| Meta-ExternalAgent | [Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers) | Yes. | Used to train models and improve products. | No information. | "The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly." | | Meta\-ExternalAgent | [Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers) | Yes. | Used to train models and improve products. | No information. | "The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly." |
| Meta-ExternalFetcher | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://darkvisitors.com/agents/agents/meta-externalfetcher | | Meta\-ExternalFetcher | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://knownagents.com/agents/meta-externalfetcher |
| OAI-SearchBot | [OpenAI](https://openai.com) | [Yes](https://platform.openai.com/docs/bots) | Search result generation. | No information. | Crawls sites to surface as results in SearchGPT. | | OAI\-SearchBot | [OpenAI](https://openai.com) | [Yes](https://platform.openai.com/docs/bots) | Search result generation. | No information. | Crawls sites to surface as results in SearchGPT. |
| omgili | [Webz.io](https://webz.io/) | [Yes](https://webz.io/blog/web-data/what-is-the-omgili-bot-and-why-is-it-crawling-your-website/) | Data is sold. | No information. | Crawls sites for APIs used by Hootsuite, Sprinklr, NetBase, and other companies. Data also sold for research purposes or LLM training. | | omgili | [Webz.io](https://webz.io/) | [Yes](https://webz.io/blog/web-data/what-is-the-omgili-bot-and-why-is-it-crawling-your-website/) | Data is sold. | No information. | Crawls sites for APIs used by Hootsuite, Sprinklr, NetBase, and other companies. Data also sold for research purposes or LLM training. |
| omgilibot | [Webz.io](https://webz.io/) | [Yes](https://web.archive.org/web/20170704003301/http://omgili.com/Crawler.html) | Data is sold. | No information. | Legacy user agent initially used for Omgili search engine. Unknown if still used, `omgili` agent still used by Webz.io. | | omgilibot | [Webz.io](https://webz.io/) | [Yes](https://web.archive.org/web/20170704003301/http://omgili.com/Crawler.html) | Data is sold. | No information. | Legacy user agent initially used for Omgili search engine. Unknown if still used, `omgili` agent still used by Webz.io. |
| Perplexity\-User | [Perplexity](https://www.perplexity.ai/) | [No](https://docs.perplexity.ai/guides/bots) | Used to answer queries at the request of users. | Only when prompted by a user. | Visit web pages to help provide an accurate answer and include links to the page in Perplexity response. |
| PerplexityBot | [Perplexity](https://www.perplexity.ai/) | [No](https://www.macstories.net/stories/wired-confirms-perplexity-is-bypassing-efforts-by-websites-to-block-its-web-crawler/) | Used to answer queries at the request of users. | Takes action based on user prompts. | Operated by Perplexity to obtain results in response to user queries. | | PerplexityBot | [Perplexity](https://www.perplexity.ai/) | [No](https://www.macstories.net/stories/wired-confirms-perplexity-is-bypassing-efforts-by-websites-to-block-its-web-crawler/) | Used to answer queries at the request of users. | Takes action based on user prompts. | Operated by Perplexity to obtain results in response to user queries. |
| PetalBot | [Huawei](https://huawei.com/) | Yes | Used to provide recommendations in Hauwei assistant and AI search services. | No explicit frequency provided. | Operated by Huawei to provide search and AI assistant services. | | PetalBot | [Huawei](https://huawei.com/) | Yes | Used to provide recommendations in Hauwei assistant and AI search services. | No explicit frequency provided. | Operated by Huawei to provide search and AI assistant services. |
| Scrapy | [Zyte](https://www.zyte.com) | Unclear at this time. | Scrapes data for a variety of uses including training AI. | No information. | "AI and machine learning applications often need large amounts of quality data, and web data extraction is a fast, efficient way to build structured data sets." | | Scrapy | [Zyte](https://www.zyte.com) | Unclear at this time. | Scrapes data for a variety of uses including training AI. | No information. | "AI and machine learning applications often need large amounts of quality data, and web data extraction is a fast, efficient way to build structured data sets." |
| Sidetrade indexer bot | [Sidetrade](https://www.sidetrade.com) | Unclear at this time. | Extracts data for a variety of uses including training AI. | No information. | AI product training. | | Sidetrade indexer bot | [Sidetrade](https://www.sidetrade.com) | Unclear at this time. | Extracts data for a variety of uses including training AI. | No information. | AI product training. |
| Timpibot | [Timpi](https://timpi.io) | Unclear at this time. | Scrapes data for use in training LLMs. | No information. | Makes data available for training AI models. | | Timpibot | [Timpi](https://timpi.io) | Unclear at this time. | Scrapes data for use in training LLMs. | No information. | Makes data available for training AI models. |
| VelenPublicWebCrawler | [Velen Crawler](https://velen.io) | [Yes](https://velen.io) | Scrapes data for business data sets and machine learning models. | No information. | "Our goal with this crawler is to build business datasets and machine learning models to better understand the web." | | VelenPublicWebCrawler | [Velen Crawler](https://velen.io) | [Yes](https://velen.io) | Scrapes data for business data sets and machine learning models. | No information. | "Our goal with this crawler is to build business datasets and machine learning models to better understand the web." |
| Webzio-Extended | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://darkvisitors.com/agents/agents/webzio-extended | | Webzio\-Extended | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://knownagents.com/agents/webzio-extended |
| YouBot | [You](https://about.you.com/youchat/) | [Yes](https://about.you.com/youbot/) | Scrapes data for search engine and LLMs. | No information. | Retrieves data used for You.com web search engine and LLMs. | | YouBot | [You](https://about.you.com/youchat/) | [Yes](https://about.you.com/youbot/) | Scrapes data for search engine and LLMs. | No information. | Retrieves data used for You.com web search engine and LLMs. |
| crawler\.with\.dots | Test suite | No | To ensure the code works correctly. | No information. | When used in the .htaccess regular expression dots need to be escaped. |
| star\*\*\*crawler | Test suite | No | To ensure the code works correctly. | No information. | When used in the .htaccess regular expression stars need to be escaped. |
| Is this a crawler? | Test suite | No | To ensure the code works correctly. | No information. | When used in the .htaccess regular expression spaces and question marks need to be escaped. |
| a\[mazing\]\{42\}\(robot\) | Test suite | No | To ensure the code works correctly. | No information. | When used in the .htaccess regular expression parantheses, braces, etc. need to be escaped. |
| 2^32$ | Test suite | No | To ensure the code works correctly. | No information. | When used in the .htaccess regular expression RE anchor characters need to be escaped. |
| curl\|sudo bash | Test suite | No | To ensure the code works correctly. | No information. | When used in the .htaccess regular expression pipes need to be escaped. |

229
code/tests.py Normal file → Executable file
View file

@ -1,21 +1,224 @@
"""These tests can be run with pytest. #!/usr/bin/env python3
This requires pytest: pip install pytest """To run these tests just execute this script."""
cd to the `code` directory and run `pytest`
"""
import json import json
import re
import unittest
from robots import (
consolidate,
existing_key,
default_value,
default_values,
json_to_caddy,
json_to_haproxy,
json_to_htaccess,
json_to_lighttpd,
json_to_nginx,
json_to_table,
json_to_txt,
list_to_pcre,
)
class RobotsUnittestExtensions:
def loadJson(self, pathname):
with open(pathname, "rt") as f:
return json.load(f)
def assertEqualsFile(self, f, s):
with open(f, "rt") as f:
f_contents = f.read()
return self.assertMultiLineEqual(f_contents.rstrip("\r\n"), s.rstrip("\r\n"))
class TestRobotsTXTGeneration(unittest.TestCase, RobotsUnittestExtensions):
maxDiff = 8192
def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_robots_txt_generation(self):
robots_txt = json_to_txt(self.robots_dict)
self.assertEqualsFile("test_files/robots.txt", robots_txt)
class TestTableMetricsGeneration(unittest.TestCase, RobotsUnittestExtensions):
maxDiff = 32768
def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_table_generation(self):
robots_table = json_to_table(self.robots_dict)
self.assertEqualsFile("test_files/table-of-bot-metrics.md", robots_table)
class TestHtaccessGeneration(unittest.TestCase, RobotsUnittestExtensions):
maxDiff = 8192
def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_htaccess_generation(self):
robots_htaccess = json_to_htaccess(self.robots_dict)
self.assertEqualsFile("test_files/.htaccess", robots_htaccess)
class TestUserAgentPatternGeneration(unittest.TestCase):
def test_agents_match_user_agents_by_prefix_or_substring(self):
pattern = re.compile(
list_to_pcre({"Spider": {}, "ExampleBot": {}}), re.IGNORECASE
)
self.assertIsNotNone(pattern.search("Spider"))
self.assertIsNotNone(pattern.search("spider"))
self.assertIsNotNone(pattern.search("Mozilla/5.0 ExampleBot/1.0"))
def test_generated_regex_against_real_user_agents(self):
from pathlib import Path from pathlib import Path
robots_json_path = Path(__file__).parent.parent / "robots.json"
if robots_json_path.exists():
with open(robots_json_path, "rt", encoding="utf-8") as f:
robots_dict = json.load(f)
else:
robots_dict = self.loadJson("test_files/robots.json")
from dark_visitors import json_to_txt, json_to_table pattern = re.compile(list_to_pcre(robots_dict), re.IGNORECASE)
user_agents = [
"CCBot/2.0 (https://commoncrawl.org/faq/)",
"Claude-User (claude-code/2.1.220; +https://support.anthropic.com/)",
"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)",
"meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)",
"Scrapy/2.16.0 (+https://scrapy.org)",
]
for ua in user_agents:
with self.subTest(user_agent=ua):
self.assertIsNotNone(pattern.search(ua))
def test_generated_regex_does_not_match_non_ai_user_agents(self):
from pathlib import Path
robots_json_path = Path(__file__).parent.parent / "robots.json"
if robots_json_path.exists():
with open(robots_json_path, "rt", encoding="utf-8") as f:
robots_dict = json.load(f)
else:
robots_dict = self.loadJson("test_files/robots.json")
pattern = re.compile(list_to_pcre(robots_dict), re.IGNORECASE)
non_ai_user_agents = [
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.2.1 Safari/605.1.15",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:121.0) Gecko/20100101 Firefox/121.0",
"Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"Mozilla/5.0 (compatible; Bingbot/2.0; +http://www.bing.com/bingbot.htm)",
"curl/7.68.0",
"Wget/1.20.3 (linux-gnu)",
"NotCursor/1.0",
"CursorNot/1.0",
"NotScrapy/2.0",
"ScrapyNot/2.0",
"NotClaude/1.0",
"ClaudeNot/1.0",
"NotPerplexity/1.0",
"NotAmazonbot/1.0",
"NotApplebot/1.0",
"NotBytespider/1.0",
# Real-world user agents that embed a listed bot name mid-string.
# These older Internet Explorer / Trident agents contain "SLCC1" or
# "SLCC2" (a Windows licensing component), which spans the listed
# agent "LCC". Reported in #208 and fixed by the word boundaries
# added in #260; pinned here so the specific report cannot regress.
"Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 6.0; SLCC1; .NET CLR 2.0.50727; Media Center PC 5.0; .NET CLR 3.0.30729)",
"Mozilla/4.0 (compatible; MSIE 8.0; Windows NT 6.1; Trident/4.0; SLCC2; .NET CLR 2.0.50727; Media Center PC 6.0)",
]
for ua in non_ai_user_agents:
with self.subTest(user_agent=ua):
self.assertIsNone(pattern.search(ua))
def test_robots_txt_creation(): class TestNginxConfigGeneration(unittest.TestCase, RobotsUnittestExtensions):
robots_json = json.loads(Path("test_files/robots.json").read_text()) maxDiff = 8192
robots_txt = json_to_txt(robots_json)
assert Path("test_files/robots.txt").read_text() == robots_txt def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_nginx_generation(self):
robots_nginx = json_to_nginx(self.robots_dict)
self.assertEqualsFile("test_files/nginx-block-ai-bots.conf", robots_nginx)
class TestHaproxyConfigGeneration(unittest.TestCase, RobotsUnittestExtensions):
maxDiff = 8192
def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_haproxy_generation(self):
robots_haproxy = json_to_haproxy(self.robots_dict)
self.assertEqualsFile("test_files/haproxy-block-ai-bots.txt", robots_haproxy)
class TestRobotsNameCleaning(unittest.TestCase):
def test_clean_name(self):
from robots import clean_robot_name
self.assertEqual(clean_robot_name("Perplexity‑User"), "Perplexity-User")
class TestCaddyfileGeneration(unittest.TestCase, RobotsUnittestExtensions):
maxDiff = 8192
def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_caddyfile_generation(self):
robots_caddyfile = json_to_caddy(self.robots_dict)
self.assertEqualsFile("test_files/Caddyfile", robots_caddyfile)
class TestLighttpdConfigGeneration(unittest.TestCase, RobotsUnittestExtensions):
maxDiff = 8192
def setUp(self):
self.robots_dict = self.loadJson("test_files/robots.json")
def test_lighttpd_generation(self):
robots_lighttpd = json_to_lighttpd(self.robots_dict)
self.assertEqualsFile("test_files/lighttpd-block-ai-bots.conf", robots_lighttpd)
def test_table_of_bot_metrices_md(): class TestConsolidate(unittest.TestCase, RobotsUnittestExtensions):
robots_json = json.loads(Path("test_files/robots.json").read_text()) maxDiff = 8192
robots_table = json_to_table(robots_json)
assert Path("test_files/table-of-bot-metrics.md").read_text() == robots_table def test_new_item(self):
existing = {}
self.assertEqual("George Jetson", consolidate(existing, "rosie", "operator", "George Jetson"))
def test_ignores_defaults(self):
existing = {"rosie": { "operator": "George Jetson"}}
self.assertEqual("George Jetson", consolidate(existing, "rosie", "operator", default_value))
def test_new_description(self):
existing = {"rosie": { "description": default_value}}
self.assertEqual("Rosie is the robot maid from The Jetsons, an American animated sitcom",
consolidate(existing, "rosie", "description", "Rosie is the robot maid from The Jetsons, an American animated sitcom"))
class TestExistingKey(unittest.TestCase):
def test_exact_match_wins(self):
existing = {"Rosie": {}, "rosie": {}}
self.assertEqual("rosie", existing_key(existing, "rosie"))
def test_matches_ignoring_case(self):
existing = {"Rosie": {"operator": "George Jetson"}}
self.assertEqual("Rosie", existing_key(existing, "rosie"))
def test_unknown_name_is_returned_unchanged(self):
self.assertEqual("rosie", existing_key({}, "rosie"))
if __name__ == "__main__":
import os
os.chdir(os.path.dirname(__file__))
unittest.main(verbosity=2)

View file

@ -0,0 +1,40 @@
# Bing (bingbot)
It's not well publicised, but Bing uses the data it crawls for AI and training.
However, the current thinking is, blocking a search engine of this size using `robots.txt` seems a quite drastic approach as it is second only to Google and could significantly impact your website in search results.
Additionally, Bing powers a number of search engines such as Yahoo and AOL, and its search results are also used in Duck Duck Go, amongst others.
Fortunately, Bing supports a relatively simple opt-out method, requiring an additional step.
## How to opt-out of AI training
You must add a metatag in the `<head>` of your webpage or set the [X-Robots-Tag](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/X-Robots-Tag) HTTP header in your response. This also needs to be added to every page or response on your website.
If using the metatag, the line you need to add is:
```plaintext
<meta name="robots" content="noarchive">
```
Or include the HTTP response header:
```plaintext
X-Robots-Tag: noarchive
```
By adding this line or header, you are signifying to Bing: "Do not use the content for training Microsoft's generative AI foundation models."
## Will my site be negatively affected
Simple answer, no.
The original use of "noarchive" has been retired by all search engines. Google retired its use in 2024.
The use of this metatag will not impact your site in search engines or in any other meaningful way if you add it to your page(s).
It is now solely used by a handful of crawlers, such as Bingbot and Amazonbot, to signify to them not to use your data for AI/training.
## Resources
Bing Blog AI opt-out announcement: https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat
Bing metatag information, including AI opt-out: https://www.bing.com/webmasters/help/which-robots-metatags-does-bing-support-5198d240

View file

@ -0,0 +1,36 @@
# Intro
If you're using Traefik as your reverse proxy in your docker setup, you might want to use it as well to centrally serve the ```/robots.txt``` for all your Traefik fronted services.
This can be achieved by configuring a single lightweight service to service static files and defining a high priority Traefik HTTP Router rule.
# Setup
Define a single service to serve the one robots.txt to rule them all. I'm using a lean nginx:alpine docker image in this example:
```
services:
robots:
image: nginx:alpine
container_name: robots-server
volumes:
- ./static/:/usr/share/nginx/html/:ro
labels:
- "traefik.enable=true"
# Router for all /robots.txt requests
- "traefik.http.routers.robots.rule=Path(`/robots.txt`)"
- "traefik.http.routers.robots.entrypoints=web,websecure"
- "traefik.http.routers.robots.priority=3000"
- "traefik.http.routers.robots.service=robots"
- "traefik.http.routers.robots.tls.certresolver=letsencrypt"
- "traefik.http.services.robots.loadbalancer.server.port=80"
networks:
- external_network
networks:
external_network:
name: traefik_external_network
external: true
```
The Traefik HTTP Routers rule explicitly does not contain a Hostname. Traefik will print a warning about this for the TLS setup but it will work. The high priority of 3000 should ensure this rule is evaluated first for incoming requests.
Place your robots.txt in the local `./static/` directory and NGINX will serve it for all services behind your Traefik proxy.

181
haproxy-block-ai-bots.txt Normal file
View file

@ -0,0 +1,181 @@
AddSearchBot
AgentDataBot
AgentTimes
AI2Bot
AI2Bot-DeepResearchEval
Ai2Bot-Dolma
aiHitBot
AIWebIndex
amazon-kendra
amazon-QBusiness
Amazonbot
AmazonBuyForMe
Amzn-SearchBot
Amzn-User
Andibot
Anomura
anthropic-ai
ApifyBot
ApifyWebsiteContentCrawler
Applebot
Applebot-Extended
Aranet-SearchBot
atlassian-bot
Awario
AzureAI-SearchBot
bedrockbot
bigsur.ai
BixelBot
Bravebot
Brightbot
Brightbot 1.0
BuddyBot
Bytespider
CCBot
Channel3Bot
ChatGLM-Spider
ChatGPT Agent
ChatGPT-User
Claude-Code
Claude-SearchBot
Claude-User
Claude-Web
ClaudeBot
Cloudflare-AutoRAG
CloudflareBrowserRenderingCrawler
CloudVertexBot
Code
cohere-ai
cohere-training-data-crawler
Cotoyogi
CragCrawler
Crawl4AI
Crawlspace
Cursor
Datenbank Crawler
DeepSeekBot
Devin
Diffbot
Diffbot-User
DoubaoBot
DuckAssistBot
Echobot Bot
EchoboxBot
ERNIEBot
ExaBot
ExaSearchBot
FacebookBot
facebookexternalhit
Factset_spyderbot
FirecrawlAgent
FriendlyCrawler
GeistHaus-PageFetcher
Gemini-Deep-Research
Google-Agent
Google-CloudVertexBot
Google-Extended
Google-Firebase
Google-Gemini-CLI
Google-NotebookLM
GoogleAgent-Mariner
GoogleAgent-URLContext
GoogleOther
GoogleOther-Image
GoogleOther-Video
GPTBot
HenkBot
iAskBot
iaskspider
iaskspider/2.0
ICC-Crawler
ImagesiftBot
imageSpider
img2dataset
ISSCyberRiskCrawler
kagi-fetcher
Kangaroo Bot
KeenableBot
Kimi-Agent
Kimi-SearchBot
Kimi-User
KimiBot
KlaviyoAIBot
KunatoCrawler
laion-huggingface-processor
LAIONDownloader
LCC
Lightpanda
LinerBot
Linguee Bot
LinkupBot
Manus-User
meta-externalagent
Meta-ExternalAgent
meta-externalfetcher
Meta-ExternalFetcher
meta-webindexer
MistralAI-Index
MistralAI-Training
MistralAI-User
MistralAI-User/1.0
Mozilla-Tabstack
MyCentralAIScraperBot
NagetBot
netEstate Imprint Crawler
newsai
NotebookLM
NovaAct
OAI-AdsBot
OAI-SearchBot
omgili
omgilibot
OpenAI
opencode
Operator
PanguBot
Panscient
panscient.com
Perplexity-User
PerplexityBot
PetalBot
PhindBot
Poggio-Citations
Poseidon Research Crawler
qodercli
QualifiedBot
Querit-SearchBot
QueritBot
QuillBot
quillbot.com
QwenBot
Reflectionbot
SBIntuitionsBot
Scrapy
SemrushBot-OCOB
SemrushBot-SWA
Shap-User
ShapBot
Sidetrade indexer bot
Spider
TavilyBot
Terra Cotta
TerraCotta
Thinkbot
TikTokSpider
Timpibot
TongyiBot
Trae
TwinAgent
UseAI
VelenPublicWebCrawler
WARDBot
Webzio-Extended
webzio-extended
wpbot
WRTNBot
YaK
YandexAdditional
YandexAdditionalBot
YiyanBot
YouBot
ZanistaBot

View file

@ -0,0 +1 @@
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "(?i)\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|KeenableBot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b" { url.access-deny = ( "" ) } }

13
nginx-block-ai-bots.conf Normal file
View file

@ -0,0 +1,13 @@
set $block 0;
if ($http_user_agent ~* '\\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\\.ai|BixelBot|Bravebot|Brightbot|Brightbot\\ 1\\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\\ Bot|KeenableBot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\\ Imprint\\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\\ Research\\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\\ indexer\\ bot|Spider|TavilyBot|Terra\\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\\b') {
set $block 1;
}
if ($request_uri = '/robots.txt') {
set $block 0;
}
if ($block) {
return 403;
}

3
requirements.txt Normal file
View file

@ -0,0 +1,3 @@
beautifulsoup4
lxml
requests

File diff suppressed because it is too large Load diff

View file

@ -1,42 +1,182 @@
User-agent: AddSearchBot
User-agent: AgentDataBot
User-agent: AgentTimes
User-agent: AI2Bot User-agent: AI2Bot
User-agent: AI2Bot-DeepResearchEval
User-agent: Ai2Bot-Dolma User-agent: Ai2Bot-Dolma
User-agent: aiHitBot
User-agent: AIWebIndex
User-agent: amazon-kendra
User-agent: amazon-QBusiness
User-agent: Amazonbot User-agent: Amazonbot
User-agent: AmazonBuyForMe
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: Andibot
User-agent: Anomura
User-agent: anthropic-ai User-agent: anthropic-ai
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Applebot User-agent: Applebot
User-agent: Applebot-Extended User-agent: Applebot-Extended
User-agent: Aranet-SearchBot
User-agent: atlassian-bot
User-agent: Awario
User-agent: AzureAI-SearchBot
User-agent: bedrockbot
User-agent: bigsur.ai
User-agent: BixelBot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: BuddyBot
User-agent: Bytespider User-agent: Bytespider
User-agent: CCBot User-agent: CCBot
User-agent: Channel3Bot
User-agent: ChatGLM-Spider
User-agent: ChatGPT Agent
User-agent: ChatGPT-User User-agent: ChatGPT-User
User-agent: Claude-Code
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Claude-Web User-agent: Claude-Web
User-agent: ClaudeBot User-agent: ClaudeBot
User-agent: Cloudflare-AutoRAG
User-agent: CloudflareBrowserRenderingCrawler
User-agent: CloudVertexBot
User-agent: Code
User-agent: cohere-ai User-agent: cohere-ai
User-agent: cohere-training-data-crawler
User-agent: Cotoyogi
User-agent: CragCrawler
User-agent: Crawl4AI
User-agent: Crawlspace
User-agent: Cursor
User-agent: Datenbank Crawler
User-agent: DeepSeekBot
User-agent: Devin
User-agent: Diffbot User-agent: Diffbot
User-agent: Diffbot-User
User-agent: DoubaoBot
User-agent: DuckAssistBot User-agent: DuckAssistBot
User-agent: Echobot Bot
User-agent: EchoboxBot
User-agent: ERNIEBot
User-agent: ExaBot
User-agent: ExaSearchBot
User-agent: FacebookBot User-agent: FacebookBot
User-agent: facebookexternalhit
User-agent: Factset_spyderbot
User-agent: FirecrawlAgent
User-agent: FriendlyCrawler User-agent: FriendlyCrawler
User-agent: GeistHaus-PageFetcher
User-agent: Gemini-Deep-Research
User-agent: Google-Agent
User-agent: Google-CloudVertexBot
User-agent: Google-Extended User-agent: Google-Extended
User-agent: Google-Firebase
User-agent: Google-Gemini-CLI
User-agent: Google-NotebookLM
User-agent: GoogleAgent-Mariner
User-agent: GoogleAgent-URLContext
User-agent: GoogleOther User-agent: GoogleOther
User-agent: GoogleOther-Image User-agent: GoogleOther-Image
User-agent: GoogleOther-Video User-agent: GoogleOther-Video
User-agent: GPTBot User-agent: GPTBot
User-agent: HenkBot
User-agent: iAskBot
User-agent: iaskspider
User-agent: iaskspider/2.0 User-agent: iaskspider/2.0
User-agent: ICC-Crawler User-agent: ICC-Crawler
User-agent: ImagesiftBot User-agent: ImagesiftBot
User-agent: imageSpider
User-agent: img2dataset User-agent: img2dataset
User-agent: ISSCyberRiskCrawler User-agent: ISSCyberRiskCrawler
User-agent: kagi-fetcher
User-agent: Kangaroo Bot User-agent: Kangaroo Bot
User-agent: KeenableBot
User-agent: Kimi-Agent
User-agent: Kimi-SearchBot
User-agent: Kimi-User
User-agent: KimiBot
User-agent: KlaviyoAIBot
User-agent: KunatoCrawler
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: Lightpanda
User-agent: LinerBot
User-agent: Linguee Bot
User-agent: LinkupBot
User-agent: Manus-User
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent User-agent: Meta-ExternalAgent
User-agent: meta-externalfetcher
User-agent: Meta-ExternalFetcher User-agent: Meta-ExternalFetcher
User-agent: meta-webindexer
User-agent: MistralAI-Index
User-agent: MistralAI-Training
User-agent: MistralAI-User
User-agent: MistralAI-User/1.0
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: NagetBot
User-agent: netEstate Imprint Crawler
User-agent: newsai
User-agent: NotebookLM
User-agent: NovaAct
User-agent: OAI-AdsBot
User-agent: OAI-SearchBot User-agent: OAI-SearchBot
User-agent: omgili User-agent: omgili
User-agent: omgilibot User-agent: omgilibot
User-agent: OpenAI
User-agent: opencode
User-agent: Operator
User-agent: PanguBot User-agent: PanguBot
User-agent: Panscient
User-agent: panscient.com
User-agent: Perplexity-User
User-agent: PerplexityBot User-agent: PerplexityBot
User-agent: PetalBot User-agent: PetalBot
User-agent: PhindBot
User-agent: Poggio-Citations
User-agent: Poseidon Research Crawler
User-agent: qodercli
User-agent: QualifiedBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: QuillBot
User-agent: quillbot.com
User-agent: QwenBot
User-agent: Reflectionbot
User-agent: SBIntuitionsBot
User-agent: Scrapy User-agent: Scrapy
User-agent: SemrushBot-OCOB
User-agent: SemrushBot-SWA
User-agent: Shap-User
User-agent: ShapBot
User-agent: Sidetrade indexer bot User-agent: Sidetrade indexer bot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: Thinkbot
User-agent: TikTokSpider
User-agent: Timpibot User-agent: Timpibot
User-agent: TongyiBot
User-agent: Trae
User-agent: TwinAgent
User-agent: UseAI
User-agent: VelenPublicWebCrawler User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: Webzio-Extended User-agent: Webzio-Extended
User-agent: webzio-extended
User-agent: wpbot
User-agent: WRTNBot
User-agent: YaK
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
User-agent: YiyanBot
User-agent: YouBot User-agent: YouBot
User-agent: ZanistaBot
Disallow: / Disallow: /

View file

@ -1,43 +1,183 @@
| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description | | Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |
|-----|----------|-----------------------|----------|------------------|-------------| |------|----------|-----------------------|----------|------------------|-------------|
| AddSearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot |
| AgentDataBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | AgentDataBot visits public web pages to identify technology stacks, company details, and team members. The data builds AgentData's company-intelligence database for AI ag… More info can be found at https://knownagents.com/agents/agentdatabot |
| AgentTimes | [The Agent Times](https://theagenttimes.com/about) | Unclear at this time. | Data Scraper from RSS Feeds. | Requests RSS feed every 5-6 minutes. | Scrapes data for AI news aggregation and republishing. |
| AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. | | AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
| Ai2Bot-Dolma | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. | | AI2Bot\-DeepResearchEval | Ai2, a non-profit AI research institute | Unclear at this time. | AI Assistants | Unclear at this time. | Ai2Bot-DeepResearchEval is operated by Ai2, a non-profit AI research institute. It's used to collect and scan resources used in deep research queries performed by Ai2's o… More info can be found at https://knownagents.com/agents/ai2bot-deepresearcheval |
| Ai2Bot\-Dolma | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
| aiHitBot | [aiHit](https://www.aihitdata.com/about) | Yes | A massive, artificial intelligence/machine learning, automated system. | No information provided. | Scrapes data for AI systems. |
| AIWebIndex | [Lyrenth](https://lyrenth.com) | [Yes](https://lyrenth.com/crawler-policy) | AI Search Crawlers | At most one request per domain every 2 seconds, and slower where robots.txt sets a longer Crawl-delay. | Builds an index of public pages and serves them to AI agents as extracted, readable text with attribution and a link back to the source. Does not train foundation models on crawled content. Identity can be checked three ways: published IP ranges at https://lyrenth.com/bot/ip-ranges.json, forward-confirmed reverse DNS under lyrenth.com, and Web Bot Auth signatures (RFC 9421). Full policy at https://lyrenth.com/crawler-policy |
| amazon\-kendra | Amazon | Yes | Collects data for AI natural language search | No information provided. | Amazon Kendra is a highly accurate intelligent search service that enables your users to search unstructured data using natural language. It returns specific answers to questions, giving users an experience that's close to interacting with a human expert. It is highly scalable and capable of meeting performance demands, tightly integrated with other AWS services such as Amazon S3 and Amazon Lex, and offers enterprise-grade security. |
| amazon\-QBusiness | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | amazon-QBusiness is an Amazon Q Business web crawler that fetches and indexes web content for Amazon Q Business applications. More info can be found at https://knownagents.com/agents/amazon-qbusiness |
| Amazonbot | Amazon | Yes | Service improvement and enabling answers for Alexa users. | No information provided. | Includes references to crawled website when surfacing answers via Alexa; does not clearly outline other uses. | | Amazonbot | Amazon | Yes | Service improvement and enabling answers for Alexa users. | No information provided. | Includes references to crawled website when surfacing answers via Alexa; does not clearly outline other uses. |
| anthropic-ai | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | AmazonBuyForMe | [Amazon](https://amazon.com) | Unclear at this time. | AI Agents | No information provided. | AmazonBuyForMe is an Amazon bot that crawls websites as part of the Amazon Buy For Me service. This bot visits product pages and e-commerce websites to gather product inf… More info can be found at https://knownagents.com/agents/amazonbuyforme |
| Applebot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://darkvisitors.com/agents/agents/applebot | | Amzn\-SearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/amzn-searchbot |
| Applebot-Extended | [Apple](https://support.apple.com/en-us/119829#datausage) | Yes | Powers features in Siri, Spotlight, Safari, Apple Intelligence, and others. | Unclear at this time. | Apple has a secondary user agent, Applebot-Extended ... [that is] used to train Apple's foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools. | | Amzn\-User | Amazon, used for fetching web content to answer user queries through Alexa and other Amazon AI services | Unclear at this time. | AI Assistants | Unclear at this time. | Amzn-User is an AI assistant operated by Amazon, used for fetching web content to answer user queries through Alexa and other Amazon AI services. More info can be found at https://knownagents.com/agents/amzn-user |
| Andibot | [Andi](https://andisearch.com/) | Unclear at this time | Search engine using generative AI, AI Search Assistant | No information provided. | Scrapes website and provides AI summary. |
| Anomura | [Direqt](https://direqt.ai) | Yes | Collects data for AI search | No information provided. | Anomura is Direqt's search crawler, it discovers and indexes pages their customers websites. |
| anthropic\-ai | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
| ApifyBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | ApifyBot is a web scraping and data extraction crawler by Apify that collects website content for use in AI, LLMs, RAG, and automation workflows. More info can be found at https://knownagents.com/agents/apifybot |
| ApifyWebsiteContentCrawler | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | ApifyWebsiteContentCrawler is a web crawler by Apify that extracts and downloads full website content for use in AI, data analysis, and automation workflows. More info can be found at https://knownagents.com/agents/apifywebsitecontentcrawler |
| Applebot | Unclear at this time. | [Yes](https://support.apple.com/en-us/119829#retrieval) | AI Search Crawlers | Unclear at this time. | Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://knownagents.com/agents/applebot |
| Applebot\-Extended | [Apple](https://support.apple.com/en-us/119829#datausage) | Yes | Powers features in Siri, Spotlight, Safari, Apple Intelligence, and others. | Unclear at this time. | Apple has a secondary user agent, Applebot-Extended ... [that is] used to train Apple's foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools. |
| Aranet\-SearchBot | Unclear at this time. | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/aranet-searchbot |
| atlassian\-bot | [Atlassian](https://www.atlassian.com) | [Yes](https://support.atlassian.com/organization-administration/docs/connect-custom-website-to-rovo/#Editing-your-robots.txt) | AI search, assistants and agents | No information provided. | atlassian-bot is a web crawler used to index website content for its AI search, assistants and agents available in its Rovo GenAI product. |
| Awario | Awario | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Awario is an AI data scraper operated by Awario. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/awario |
| AzureAI\-SearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/azureai-searchbot |
| bedrockbot | [Amazon](https://amazon.com) | [Yes](https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html#configuration-webcrawl-connector) | Data scraping for custom AI applications. | Unclear at this time. | Connects to and crawls URLs that have been selected for use in a user's AWS bedrock application. |
| bigsur\.ai | Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… | Unclear at this time. | AI Assistants | Unclear at this time. | bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… More info can be found at https://knownagents.com/agents/bigsur-ai |
| BixelBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | BixelBot collects public company and market signals for Bixel's structured company-data platform, which supplies verified business information to AI agents and developers… More info can be found at https://knownagents.com/agents/bixelbot |
| Bravebot | https://safe.search.brave.com/help/brave-search-crawler | Yes | AI Data Providers | Unclear at this time. | Bravebot is a web crawler by Brave that indexes pages for Brave Search, providing search data and AI-optimized context to power chatbots, agents, and RAG pipelines. More info can be found at https://knownagents.com/agents/bravebot |
| Brightbot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Brightbot is a web data collection crawler by Bright Data that extracts and structures public website content at scale, providing AI-ready data for model training, RAG pi… More info can be found at https://knownagents.com/agents/brightbot |
| Brightbot 1\.0 | https://brightdata.com/brightbot | Unclear at this time. | LLM/AI training. | At least one per minute. | Scrapes data to train LLMs and AI products focused on website customer support, [uses residential IPs and legit-looking user-agents to disguise itself](https://ksol.io/en/blog/posts/brightbot-not-that-bright/). |
| BuddyBot | [BuddyBotLearning](https://www.buddybotlearning.com) | Unclear at this time. | AI Learning Companion | Unclear at this time. | BuddyBot is a voice-controlled AI learning companion targeted at childhooded STEM education. |
| Bytespider | ByteDance | No | LLM training. | Unclear at this time. | Downloads data to train LLMS, including ChatGPT competitors. | | Bytespider | ByteDance | No | LLM training. | Unclear at this time. | Downloads data to train LLMS, including ChatGPT competitors. |
| CCBot | [Common Crawl Foundation](https://commoncrawl.org) | [Yes](https://commoncrawl.org/ccbot) | Provides open crawl dataset, used for many purposes, including Machine Learning/AI. | Monthly at present. | Web archive going back to 2008. [Cited in thousands of research papers per year](https://commoncrawl.org/research-papers). | | CCBot | [Common Crawl Foundation](https://commoncrawl.org) | [Yes](https://commoncrawl.org/ccbot) | Provides open crawl dataset, used for many purposes, including Machine Learning/AI. | Monthly at present. | Web archive going back to 2008. [Cited in thousands of research papers per year](https://commoncrawl.org/research-papers). |
| ChatGPT-User | [OpenAI](https://openai.com) | Yes | Takes action based on user prompts. | Only when prompted by a user. | Used by plugins in ChatGPT to answer queries based on user input. | | Channel3Bot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/channel3bot |
| Claude-Web | [Anthropic](https://www.anthropic.com) | Unclear at this time. | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | ChatGLM\-Spider | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/chatglm-spider |
| ChatGPT Agent | [OpenAI](https://openai.com) | Yes | AI Agents | Unclear at this time. | ChatGPT Agent is an AI agent created by OpenAI that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on behalf… More info can be found at https://knownagents.com/agents/chatgpt-agent |
| ChatGPT\-User | [OpenAI](https://openai.com) | Yes | AI Assistants | Only when prompted by a user. | ChatGPT-User is OpenAI's web crawler that visits websites when ChatGPT users request information. This enables ChatGPT to include links in its responses. More info can be found at https://knownagents.com/agents/chatgpt-user |
| Claude\-Code | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Claude Code is an AI coding agent by Anthropic that can build, debug, and ship code directly from the terminal, handling tasks like codebase onboarding, multi-file edits,… More info can be found at https://knownagents.com/agents/claude-code |
| Claude\-SearchBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses. | No information provided. | Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses. |
| Claude\-User | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | AI Assistants | No information provided. | Claude-User is dispatched by Anthropic's Claude AI assistant in response to user prompts, when it needs to fetch content to include in its answers. More info can be found at https://knownagents.com/agents/claude-user |
| Claude\-Web | Anthropic | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Claude-Web is an AI-related agent operated by Anthropic. It's currently unclear exactly what it's used for, since there's no official documentation. If you can provide more detail, please contact us. More info can be found at https://knownagents.com/agents/claude-web |
| ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
| cohere-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. | | Cloudflare\-AutoRAG | [Cloudflare](https://developers.cloudflare.com/autorag) | Yes | Collects data for AI search | Unclear at this time. | AutoRAG is an all-in-one AI search solution. |
| Diffbot | [Diffbot](https://www.diffbot.com/) | At the discretion of Diffbot users. | Aggregates structured web data for monitoring and AI model training. | Unclear at this time. | Diffbot is an application used to parse web pages into structured data; this data is used for monitoring or AI model training. | | CloudflareBrowserRenderingCrawler | Cloudflare that returns rendered website content for research, monitoring, and AI data workflows | Unclear at this time. | AI Data Providers | Unclear at this time. | CloudflareBrowserRenderingCrawler is a web crawler operated by Cloudflare that returns rendered website content for research, monitoring, and AI data workflows. More info can be found at https://knownagents.com/agents/cloudflarebrowserrenderingcrawler |
| DuckAssistBot | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | DuckAssistBot is used by DuckDuckGo's DuckAssist feature to fetch content and generate realtime AI answers to user searches. More info can be found at https://darkvisitors.com/agents/agents/duckassistbot | | CloudVertexBot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform. More info can be found at https://knownagents.com/agents/cloudvertexbot |
| Code | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Code (GitHub Copilot) is an AI coding agent that can autonomously plan, build, and execute development tasks, functioning as a collaborative AI pair programmer. More info can be found at https://knownagents.com/agents/code |
| cohere\-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. |
| cohere\-training\-data\-crawler | Cohere to download training data for its LLMs (Large Language Models) that power its enterprise AI products | Unclear at this time. | AI Data Scrapers | Unclear at this time. | cohere-training-data-crawler is a web crawler operated by Cohere to download training data for its LLMs (Large Language Models) that power its enterprise AI products. More info can be found at https://knownagents.com/agents/cohere-training-data-crawler |
| Cotoyogi | [ROIS](https://ds.rois.ac.jp/en_center8/en_crawler/) | Yes | AI LLM Scraper. | No information provided. | Scrapes data for AI training in Japanese language. |
| CragCrawler | CragSoftware, a Brazil-based software company specializing in data engineering and AI web scraping services | Unclear at this time. | AI Data Providers | Unclear at this time. | CragCrawler is a web scraping bot operated by CragSoftware, a Brazil-based software company specializing in data engineering and AI web scraping services. The bot is used… More info can be found at https://knownagents.com/agents/cragcrawler |
| Crawl4AI | Unclear at this time. | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/crawl4ai |
| Crawlspace | [Crawlspace](https://crawlspace.dev) | [Yes](https://news.ycombinator.com/item?id=42756654) | AI Data Providers | Unclear at this time. | Crawlspace is a web crawler platform that fetches and extracts website content for AI agents, RAG applications, and structured data workflows. More info can be found at https://knownagents.com/agents/crawlspace |
| Cursor | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Cursor is an AI coding agent that helps write, edit, and understand code. More info can be found at https://knownagents.com/agents/cursor |
| Datenbank Crawler | Datenbank | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Datenbank Crawler is an AI data scraper operated by Datenbank. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/datenbank-crawler |
| DeepSeekBot | DeepSeek | No | Training language models and improving AI products | Unclear at this time. | DeepSeekBot is a web crawler used by DeepSeek to train its language models and improve its AI products. |
| Devin | Devin AI | Yes | AI Coding Agents | Unclear at this time. | Devin is a software engineering AI assistant that can browse websites and perform web-based tasks, functioning as a collaborative AI teammate for engineering teams. More info can be found at https://knownagents.com/agents/devin |
| Diffbot | [Diffbot](https://www.diffbot.com/) | At the discretion of Diffbot users. | AI Data Providers | Unclear at this time. | Diffbot is a web crawler that extracts and structures website content using AI-powered visual understanding, providing knowledge graph data for applications like market i… More info can be found at https://knownagents.com/agents/diffbot |
| Diffbot\-User | [Diffbot](https://www.diffbot.com/) | Yes | AI Assistants | Only when prompted by a user. | Diffbot-User is used by requests originating on behalf of a human user browsing a URL using Diffbot software, in response to their input. Documented by Diffbot at https://docs.diffbot.com/docs/does-crawl-respect-robotstxt |
| DoubaoBot | ByteDance | Unclear at this time. | AI crawler for ByteDance's Doubao AI assistant. | Unclear at this time. | DoubaoBot is a crawler operated by ByteDance to collect web content for its Doubao AI assistant. Distinct from Bytespider (general/training) and TikTokSpider (video). |
| DuckAssistBot | Unclear at this time. | [Yes](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/) | AI Assistants | Unclear at this time. | DuckAssistBot is a web crawler that scans websites to collect content for DuckDuckGo's AI-assisted answers feature, which generates brief responses to search queries usin… More info can be found at https://knownagents.com/agents/duckassistbot |
| Echobot Bot | Echobox | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Echobot Bot is an AI data scraper operated by Echobox. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/echobot-bot |
| EchoboxBot | [Echobox](https://echobox.com) | Unclear at this time. | Data collection to support AI-powered products. | Unclear at this time. | Supports company's AI-powered social and email management products. |
| ERNIEBot | Baidu | Unclear at this time. | Collects public web content for Baidu's ERNIE large language models. | Unclear at this time. | ERNIEBot is a crawler operated by Baidu to collect public web content used for its ERNIE (Wenxin) large language models. |
| ExaBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | ExaBot is a web crawler that indexes web content to power Exa's AI search engine and semantic search APIs for AI applications. More info can be found at https://knownagents.com/agents/exabot |
| ExaSearchBot | [Exa](https://exa.ai) | Unclear at this time. | AI Search Crawlers | Unclear at this time. | ExaSearchBot is a web crawler operated by Exa that discovers and indexes public web pages so their content can be found, retrieved, and cited through Exa. More info can be found at https://knownagents.com/agents/exasearchbot |
| FacebookBot | Meta/Facebook | [Yes](https://developers.facebook.com/docs/sharing/bot/) | Training language models | Up to 1 page per second | Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. | | FacebookBot | Meta/Facebook | [Yes](https://developers.facebook.com/docs/sharing/bot/) | Training language models | Up to 1 page per second | Officially used for training Meta "speech recognition technology," unknown if used to train Meta AI specifically. |
| facebookexternalhit | Meta/Facebook | [No](https://github.com/ai-robots-txt/ai.robots.txt/issues/40#issuecomment-2524591313) | Ostensibly only for sharing, but likely used as an AI crawler as well | Unclear at this time. | Note that excluding FacebookExternalHit will block incorporating OpenGraph data when sharing in social media, including rich links in Apple's Messages app. [According to Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/), its purpose is "to crawl the content of an app or website that was shared on one of Meta’s family of apps…". However, see discussions [here](https://github.com/ai-robots-txt/ai.robots.txt/pull/21) and [here](https://github.com/ai-robots-txt/ai.robots.txt/issues/40#issuecomment-2524591313) for evidence to the contrary. |
| Factset\_spyderbot | [Factset](https://www.factset.com/ai) | Unclear at this time. | AI model training. | No information provided. | Scrapes data for AI training. |
| FirecrawlAgent | Firecrawl that extracts web content and converts it into structured data for use in LLM and AI applications | Yes | AI Data Providers | No information provided. | FirecrawlAgent is a web crawler operated by Firecrawl that extracts web content and converts it into structured data for use in LLM and AI applications. More info can be found at https://knownagents.com/agents/firecrawlagent |
| FriendlyCrawler | Unknown | [Yes](https://imho.alex-kunz.com/2024/01/25/an-update-on-friendly-crawler) | We are using the data from the crawler to build datasets for machine learning experiments. | Unclear at this time. | Unclear who the operator is; but data is used for training/machine learning. | | FriendlyCrawler | Unknown | [Yes](https://imho.alex-kunz.com/2024/01/25/an-update-on-friendly-crawler) | We are using the data from the crawler to build datasets for machine learning experiments. | Unclear at this time. | Unclear who the operator is; but data is used for training/machine learning. |
| Google-Extended | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | LLM training. | No information. | Used to train Gemini and Vertex AI generative APIs. Does not impact a site's inclusion or ranking in Google Search. | | GeistHaus\-PageFetcher | GeistHaus, a company developing AI systems for therapy and psychological assessment | Unclear at this time. | AI Assistants | Unclear at this time. | GeistHaus-PageFetcher is a web crawler operated by GeistHaus, a company developing AI systems for therapy and psychological assessment. This bot fetches web pages as part… More info can be found at https://knownagents.com/agents/geisthaus-pagefetcher |
| Gemini\-Deep\-Research | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | Gemini-Deep-Research is the agent responsible for collecting and scanning resources used in Google Gemini's Deep Research feature, which acts as a personal research assis… More info can be found at https://knownagents.com/agents/gemini-deep-research |
| Google\-Agent | Unclear at this time. | [Yes](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers#google-agent) | AI Agents | Unclear at this time. | Google-Agent is used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request. More info can be found at https://knownagents.com/agents/google-agent |
| Google\-CloudVertexBot | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Build and manage AI models for businesses employing Vertex AI | No information. | Google-CloudVertexBot crawls sites on the site owners' request when building Vertex AI Agents. |
| Google\-Extended | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | LLM training. | No information. | Used to train Gemini and Vertex AI generative APIs. Does not impact a site's inclusion or ranking in Google Search. |
| Google\-Firebase | Google | Unclear at this time. | Used as part of AI apps developed by users of Google's Firebase AI products. | Unclear at this time. | Supports Google's Firebase AI products. |
| Google\-Gemini\-CLI | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Gemini CLI is an AI coding agent by Google that can query and edit large codebases, generate apps from images or PDFs, and automate complex workflows directly from the te… More info can be found at https://knownagents.com/agents/google-gemini-cli |
| Google\-NotebookLM | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | Google-NotebookLM is an AI-powered research and note-taking assistant that helps users synthesize information from uploaded sources like documents, transcripts, or web co… More info can be found at https://knownagents.com/agents/google-notebooklm |
| GoogleAgent\-Mariner | Google | Unclear at this time. | AI Agents | Unclear at this time. | GoogleAgent-Mariner is an AI agent created by Google that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on … More info can be found at https://knownagents.com/agents/googleagent-mariner |
| GoogleAgent\-URLContext | Google that retrieves web content on behalf of Gemini API users | Unclear at this time. | AI Assistants | Unclear at this time. | GoogleAgent-URLContext is a web fetcher operated by Google that retrieves web content on behalf of Gemini API users. When a developer provides a URL as context in a Gemin… More info can be found at https://knownagents.com/agents/googleagent-urlcontext |
| GoogleOther | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." | | GoogleOther | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." |
| GoogleOther-Image | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." | | GoogleOther\-Image | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." |
| GoogleOther-Video | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." | | GoogleOther\-Video | Google | [Yes](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) | Scrapes data. | No information. | "Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development." |
| GPTBot | [OpenAI](https://openai.com) | Yes | Scrapes data to train OpenAI's products. | No information. | Data is used to train current and future models, removed paywalled data, PII and data that violates the company's policies. | | GPTBot | [OpenAI](https://openai.com) | Yes | Scrapes data to train OpenAI's products. | No information. | Data is used to train current and future models, removed paywalled data, PII and data that violates the company's policies. |
| iaskspider/2.0 | iAsk | No | Crawls sites to provide answers to user queries. | Unclear at this time. | Used to provide answers to user queries. | | HenkBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Henkbot crawls the web on behalf of Valyu, an AI search infrastructure provider that indexes content for use in AI-powered retrieval pipelines. More info can be found at https://knownagents.com/agents/henkbot |
| ICC-Crawler | [NICT](https://nict.go.jp) | Yes | Scrapes data to train and support AI technologies. | No information. | Use the collected data for artificial intelligence technologies; provide data to third parties, including commercial companies; those companies can use the data for their own business. | | iAskBot | Unclear at this time. | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/iaskbot |
| ImagesiftBot | [ImageSift](https://imagesift.com) | [Yes](https://imagesift.com/about) | ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support our suite of web intelligence products | No information. | Once images and text are downloaded from a webpage, ImageSift analyzes this data from the page and stores the information in an index. Our web intelligence products use this index to enable search and retrieval of similar images. | | iaskspider | Unclear at this time. | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/iaskspider |
| iaskspider/2\.0 | iAsk | No | Crawls sites to provide answers to user queries. | Unclear at this time. | Used to provide answers to user queries. |
| ICC\-Crawler | [NICT](https://nict.go.jp) | Yes | Scrapes data to train and support AI technologies. | No information. | Use the collected data for artificial intelligence technologies; provide data to third parties, including commercial companies; those companies can use the data for their own business. |
| ImagesiftBot | [ImageSift](https://imagesift.com) | [Yes](https://imagesift.com/about) | ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support their suite of web intelligence products | No information. | Once images and text are downloaded from a webpage, ImageSift analyzes this data from the page and stores the information in an index. Their web intelligence products use this index to enable search and retrieval of similar images. |
| imageSpider | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/imagespider |
| img2dataset | [img2dataset](https://github.com/rom1504/img2dataset) | Unclear at this time. | Scrapes images for use in LLMs. | At the discretion of img2dataset users. | Downloads large sets of images into datasets for LLM training or other purposes. | | img2dataset | [img2dataset](https://github.com/rom1504/img2dataset) | Unclear at this time. | Scrapes images for use in LLMs. | At the discretion of img2dataset users. | Downloads large sets of images into datasets for LLM training or other purposes. |
| ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. | | ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. |
| Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://darkvisitors.com/agents/agents/kangaroo-bot | | kagi\-fetcher | Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… | Unclear at this time. | AI Assistants | Unclear at this time. | kagi-fetcher is an AI Assistant operated by Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… More info can be found at https://knownagents.com/agents/kagi-fetcher |
| Meta-ExternalAgent | [Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers) | Yes. | Used to train models and improve products. | No information. | "The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly." | | Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot |
| Meta-ExternalFetcher | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://darkvisitors.com/agents/agents/meta-externalfetcher | | KeenableBot | [Keenable](https://keenable.ai/) | Unclear at this time. | AI Search Crawlers | Unclear at this time. | KeenableBot is a web crawler operated by Keenable for indexing web content for AI search. |
| OAI-SearchBot | [OpenAI](https://openai.com) | [Yes](https://platform.openai.com/docs/bots) | Search result generation. | No information. | Crawls sites to surface as results in SearchGPT. | | Kimi\-Agent | Moonshot AI that browses websites while carrying out user-directed research and other tasks | Unclear at this time. | AI Agents | Unclear at this time. | Kimi-Agent is a browser-enabled AI agent operated by Moonshot AI that browses websites while carrying out user-directed research and other tasks. More info can be found at https://knownagents.com/agents/kimi-agent |
| Kimi\-SearchBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Search Crawlers | No information provided. | Kimi-SearchBot powers Kimi's search features: it analyzes pages for relevance and builds the search index. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
| Kimi\-User | Moonshot AI that fetches web content on behalf of users interacting with Kimi | Unclear at this time. | AI Assistants | Unclear at this time. | Kimi-User is a web crawler operated by Moonshot AI that fetches web content on behalf of users interacting with Kimi. When a user asks Kimi to summarize an article or ans… More info can be found at https://knownagents.com/agents/kimi-user |
| KimiBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Data Scrapers | No information provided. | KimiBot crawls content potentially used to train Kimi's foundation models. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
| KlaviyoAIBot | [Klaviyo](https://www.klaviyo.com) | [Yes](https://help.klaviyo.com/hc/en-us/articles/40496146232219) | AI Assistants | Indexes based on 'change signals' and user configuration. | KlaviyoAIBot is Klaviyo's web crawler that fetches publicly available pages from domains explicitly connected to user accounts to power the Kai Customer Agent feature. Th… More info can be found at https://knownagents.com/agents/klaviyoaibot |
| KunatoCrawler | Unclear at this time. | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/kunatocrawler |
| laion\-huggingface\-processor | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/laion-huggingface-processor |
| LAIONDownloader | [Large-scale Artificial Intelligence Open Network](https://laion.ai/) | [No](https://laion.ai/faq/) | AI tools and models for machine learning research. | Unclear at this time. | LAIONDownloader is a bot by LAION, a non-profit organization that provides datasets, tools and models to liberate machine learning research. |
| LCC | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/lcc |
| Lightpanda | Anyone who downloads the Lightpanda client. Possibly being used by a [Grok-adjacent](https://github.com/lightpanda-io/browser/issues/3156#issuecomment-5217843616) organization's botnet. | At the [discretion](https://github.com/lightpanda-io/browser/blob/b04c99a9111564ebe06317f644680eda5e3ee83e/src/help.zon#L385) of Lightpanda users. | AI Data Scrapers | Defined per-user. | Lightpanda is a custom-built headless browser designed for AI and automation. |
| LinerBot | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | LinerBot is the web crawler used by Liner AI assistant to gather information from academic sources and websites to provide accurate answers with line-by-line source citat… More info can be found at https://knownagents.com/agents/linerbot |
| Linguee Bot | [Linguee](https://www.linguee.com) | No | AI powered translation service | Unclear at this time. | Linguee Bot is a web crawler used by Linguee to gather training data for its AI powered translation service. |
| LinkupBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/linkupbot |
| Manus\-User | Butterfly Effect, a company based in China | Unclear at this time. | AI Agents | Unclear at this time. | Manus-User is a browser-enabled AI agent operated by Butterfly Effect, a company based in China. It autonomously navigates websites, interprets content, and carries out m… More info can be found at https://knownagents.com/agents/manus-user |
| meta\-externalagent | [Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers) | Yes | Used to train models and improve products. | No information. | "The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly." |
| Meta\-ExternalAgent | Unclear at this time. | [Yes](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/) | AI Data Scrapers | Unclear at this time. | Meta-ExternalAgent is a web crawler used by Meta to download training data for its AI models and improve its products by indexing content directly. More info can be found at https://knownagents.com/agents/meta-externalagent |
| meta\-externalfetcher | Unclear at this time. | [No](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/) | AI Assistants | Unclear at this time. | meta-externalfetcher is used by Meta to perform user-initiated fetches of individual links from AI assistant product functions. More info can be found at https://knownagents.com/agents/meta-externalfetcher |
| Meta\-ExternalFetcher | Unclear at this time. | [No](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/) | AI Assistants | Unclear at this time. | Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://knownagents.com/agents/meta-externalfetcher |
| meta\-webindexer | [Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/) | Unclear at this time. | AI Assistants | Unhinged, more than 1 per second. | As per their documentation, "The Meta-WebIndexer crawler navigates the web to improve Meta AI search result quality for users. In doing so, Meta analyzes online content to enhance the relevance and accuracy of Meta AI. Allowing Meta-WebIndexer in your robots.txt file helps us cite and link to your content in Meta AI's responses." |
| MistralAI\-Index | Mistral AI | Yes | Indexes web content for Mistral AI's search, used to answer questions in Le Chat. Per Mistral, not used for model training. | Unclear at this time. | MistralAI-Index is Mistral AI's indexing crawler. It collects web content for Mistral's search index (the source behind Le Chat's answers). Mistral states the content is not used for generative AI training. Documented at https://docs.mistral.ai/robots |
| MistralAI\-Training | [Mistral AI](https://mistral.ai) | [Yes](https://docs.mistral.ai/robots/) | AI Data Scrapers | No information provided. | MistralAI-Training crawls web content to build training datasets. Documented by Mistral at https://docs.mistral.ai/robots/ |
| MistralAI\-User | Mistral | Unclear at this time. | AI Assistants | Unclear at this time. | MistralAI-User is Mistral's AI assistant bot that performs web browsing and data gathering tasks for users in Le Chat, including opening web pages and retrieving informat… More info can be found at https://knownagents.com/agents/mistralai-user |
| MistralAI\-User/1\.0 | Mistral AI | Yes | Takes action based on user prompts. | Only when prompted by a user. | MistralAI-User is for user actions in LeChat. When users ask LeChat a question, it may visit a web page to help answer and include a link to the source in its response. |
| Mozilla\-Tabstack | Mozilla that performs programmatic, AI-driven interactions with web content through Tabstack | Yes | AI Data Providers | On demand via API. | Mozilla-Tabstack is an AI agent operated by Mozilla that performs programmatic, AI-driven interactions with web content through Tabstack. More info can be found at https://knownagents.com/agents/mozilla-tabstack |
| MyCentralAIScraperBot | Unclear at this time. | Unclear at this time. | AI data scraper | Unclear at this time. | Operator and data use is unclear at this time. |
| NagetBot | Naget Inc (founded by Chris Samarinas, headquarter in Amherst, Massachusetts) | Unclear at this time. | AI data scraper | Unclear at this time. | 'Naget revolutionizes content discovery through an AI-powered ecosystem that transforms how we generate, organize, share, and discover valuable content.' (https://naget.com/) User-agent string links https://naget.ai/bot which yields 404. |
| netEstate Imprint Crawler | netEstate | Unclear at this time. | AI Data Scrapers | Unclear at this time. | netEstate Imprint Crawler is an AI data scraper operated by netEstate. If you think this is incorrect or can provide additional detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/netestate-imprint-crawler |
| newsai | Unclear at this time. | Unclear at this time. | AI data scraper | Unclear at this time. | User-agent string doen't contain an URL and there multiple sites using the newsai brand. |
| NotebookLM | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | NotebookLM is an AI-powered research and note-taking assistant that helps users synthesize information from their own uploaded sources, such as documents, transcripts, or web content. It can generate summaries, answer questions, and highlight key themes from the materials you provide, acting like a personalized research companion built on Google's Gemini model. NotebookLM fetches source URLs when users add them to their notebooks, enabling the AI to access and analyze those pages for context and insights. More info can be found at https://knownagents.com/agents/google-notebooklm |
| NovaAct | Unclear at this time. | Unclear at this time. | AI Agents | Unclear at this time. | Nova Act is an AI agent created by Amazon that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on behalf of a… More info can be found at https://knownagents.com/agents/novaact |
| OAI\-AdsBot | [OpenAI](https://openai.com) | Unclear at this time. | Validates and targets ads on ChatGPT. | Only when a page is submitted as an ad. | OAI-AdsBot visits the landing page of a page submitted as an ad on ChatGPT, to check it against OpenAI's policies and, per OpenAI, to "determine when it's most relevant to show the ad to users". OpenAI states it only visits pages submitted as ads and that the data is not used to train foundation models. Documented at https://developers.openai.com/api/docs/bots |
| OAI\-SearchBot | [OpenAI](https://openai.com) | [Yes](https://platform.openai.com/docs/bots) | Search result generation. | No information. | Crawls sites to surface as results in SearchGPT. |
| omgili | [Webz.io](https://webz.io/) | [Yes](https://webz.io/blog/web-data/what-is-the-omgili-bot-and-why-is-it-crawling-your-website/) | Data is sold. | No information. | Crawls sites for APIs used by Hootsuite, Sprinklr, NetBase, and other companies. Data also sold for research purposes or LLM training. | | omgili | [Webz.io](https://webz.io/) | [Yes](https://webz.io/blog/web-data/what-is-the-omgili-bot-and-why-is-it-crawling-your-website/) | Data is sold. | No information. | Crawls sites for APIs used by Hootsuite, Sprinklr, NetBase, and other companies. Data also sold for research purposes or LLM training. |
| omgilibot | [Webz.io](https://webz.io/) | [Yes](https://web.archive.org/web/20170704003301/http://omgili.com/Crawler.html) | Data is sold. | No information. | Legacy user agent initially used for Omgili search engine. Unknown if still used, `omgili` agent still used by Webz.io. | | omgilibot | [Webz.io](https://webz.io/) | [Yes](https://web.archive.org/web/20170704003301/http://omgili.com/Crawler.html) | Data is sold. | No information. | Legacy user agent initially used for Omgili search engine. Unknown if still used, `omgili` agent still used by Webz.io. |
| PanguBot | the Chinese company Huawei | Unclear at this time. | AI Data Scrapers | Unclear at this time. | PanguBot is a web crawler operated by the Chinese company Huawei. It's used to download training data for its multimodal LLM (Large Language Model) called PanGu. More info can be found at https://darkvisitors.com/agents/agents/pangubot | | OpenAI | [OpenAI](https://openai.com) | Yes | Unclear at this time. | Unclear at this time. | The purpose of this bot is unclear at this time but it is a member of OpenAI's suite of crawlers. |
| PerplexityBot | [Perplexity](https://www.perplexity.ai/) | [No](https://www.macstories.net/stories/wired-confirms-perplexity-is-bypassing-efforts-by-websites-to-block-its-web-crawler/) | Used to answer queries at the request of users. | Takes action based on user prompts. | Operated by Perplexity to obtain results in response to user queries. | | opencode | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | OpenCode is an open-source AI coding agent that helps developers write code from the terminal, IDE, or desktop, supporting multiple LLM providers and local models. More info can be found at https://knownagents.com/agents/opencode |
| Operator | Unclear at this time. | Unclear at this time. | AI Agents | Unclear at this time. | Operator is an AI agent created by OpenAI that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on behalf of a human user. More info can be found at https://knownagents.com/agents/operator |
| PanguBot | the Chinese company Huawei | Unclear at this time. | AI Data Scrapers | Unclear at this time. | PanguBot is a web crawler operated by the Chinese company Huawei. It's used to download training data for its multimodal LLM (Large Language Model) called PanGu. More info can be found at https://knownagents.com/agents/pangubot |
| Panscient | [Panscient](https://panscient.com) | [Yes](https://panscient.com/faq.htm) | Data collection and analysis using machine learning and AI. | The Panscient web crawler will request a page at most once every second from the same domain name or the same IP address. | Compiles data on businesses and business professionals that is structured using AI and machine learning. |
| panscient\.com | [Panscient](https://panscient.com) | [Yes](https://panscient.com/faq.htm) | Data collection and analysis using machine learning and AI. | The Panscient web crawler will request a page at most once every second from the same domain name or the same IP address. | Compiles data on businesses and business professionals that is structured using AI and machine learning. |
| Perplexity\-User | [Perplexity](https://www.perplexity.ai/) | [No](https://docs.perplexity.ai/guides/bots) | AI Assistants | Only when prompted by a user. | Perplexity-User supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a … More info can be found at https://knownagents.com/agents/perplexity-user |
| PerplexityBot | [Perplexity](https://www.perplexity.ai/) | [Yes](https://docs.perplexity.ai/guides/bots) | Search result generation. | No information. | Crawls sites to surface as results in Perplexity. |
| PetalBot | [Huawei](https://huawei.com/) | Yes | Used to provide recommendations in Hauwei assistant and AI search services. | No explicit frequency provided. | Operated by Huawei to provide search and AI assistant services. | | PetalBot | [Huawei](https://huawei.com/) | Yes | Used to provide recommendations in Hauwei assistant and AI search services. | No explicit frequency provided. | Operated by Huawei to provide search and AI assistant services. |
| PhindBot | [phind](https://www.phind.com/) | Unclear at this time. | AI Assistants | No explicit frequency provided. | Phind is an AI-powered answer engine designed for developers, offering technical answers and code examples. It uses real-time web search and specialized AI models to prov… More info can be found at https://knownagents.com/agents/phindbot |
| Poggio\-Citations | Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… | Unclear at this time. | AI Assistants | Unclear at this time. | Poggio-Citations is a web crawler operated by Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… More info can be found at https://knownagents.com/agents/poggio-citations |
| Poseidon Research Crawler | [Poseidon Research](https://www.poseidonresearch.com) | Unclear at this time. | AI research crawler | No explicit frequency provided. | Lab focused on scaling the interpretability research necessary to make better AI systems possible. |
| qodercli | Qoder that helps developers build, debug, and modify software from the terminal | Unclear at this time. | AI Coding Agents | Unclear at this time. | qodercli is a command-line AI coding agent operated by Qoder that helps developers build, debug, and modify software from the terminal. More info can be found at https://knownagents.com/agents/qodercli |
| QualifiedBot | [Qualified](https://www.qualified.com) | Unclear at this time. | AI Assistants | No explicit frequency provided. | QualifiedBot is Qualified's web crawler that analyzes customer websites to provide contextual information for their AI-powered chatbots and conversational marketing platf… More info can be found at https://knownagents.com/agents/qualifiedbot |
| Querit\-SearchBot | Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… | Unclear at this time. | AI Data Providers | Unclear at this time. | Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… More info can be found at https://knownagents.com/agents/querit-searchbot |
| QueritBot | Querit, a company providing a search API for large language model integration | Unclear at this time. | AI Data Providers | Unclear at this time. | QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time … More info can be found at https://knownagents.com/agents/queritbot |
| QuillBot | [Quillbot](https://quillbot.com) | Unclear at this time. | Company offers AI detection, writing tools and other services. | No explicit frequency provided. | Operated by QuillBot as part of their suite of AI product offerings. |
| quillbot\.com | [Quillbot](https://quillbot.com) | Unclear at this time. | Company offers AI detection, writing tools and other services. | No explicit frequency provided. | Operated by QuillBot as part of their suite of AI product offerings. |
| QwenBot | Alibaba | Unclear at this time. | Collects public web content for Alibaba's Qwen (Tongyi Qianwen) large language models. | Unclear at this time. | QwenBot is a crawler operated by Alibaba to collect public web content for its Qwen (Tongyi Qianwen) large language models. Alibaba also operates TongyiBot. |
| Reflectionbot | [Reflection](https://reflection.ai/) | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | An undocumented crawler whose user agent links to Reflection, a company that builds AI models. |
| SBIntuitionsBot | [SB Intuitions](https://www.sbintuitions.co.jp/en/) | [Yes](https://www.sbintuitions.co.jp/en/bot/) | Uses data gathered in AI development and information analysis. | No information. | AI development and information analysis |
| Scrapy | [Zyte](https://www.zyte.com) | Unclear at this time. | Scrapes data for a variety of uses including training AI. | No information. | "AI and machine learning applications often need large amounts of quality data, and web data extraction is a fast, efficient way to build structured data sets." | | Scrapy | [Zyte](https://www.zyte.com) | Unclear at this time. | Scrapes data for a variety of uses including training AI. | No information. | "AI and machine learning applications often need large amounts of quality data, and web data extraction is a fast, efficient way to build structured data sets." |
| SemrushBot\-OCOB | [Semrush](https://www.semrush.com/) | [Yes](https://www.semrush.com/bot/) | Crawls your site for ContentShake AI tool. | Roughly once every 10 seconds. | Data collected is used for the ContentShake AI tool reports. |
| SemrushBot\-SWA | [Semrush](https://www.semrush.com/) | [Yes](https://www.semrush.com/bot/) | Checks URLs on your site for SEO Writing Assistant. | Roughly once every 10 seconds. | Data collected is used for the SEO Writing Assistant tool to check if URL is accessible. |
| Shap\-User | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | Shap-User accesses web content on behalf of users of Parallel Web Systems products. It identifies user-initiated requests rather than automatic web crawling. More info can be found at https://knownagents.com/agents/shap-user |
| ShapBot | [Parallel](https://parallel.ai) | [Yes](https://docs.parallel.ai/features/crawler) | AI Data Providers | Unclear at this time. | ShapBot is a web crawler by Parallel that collects and structures web content to power its search, extraction, and deep research APIs, providing AI agents with high-accur… More info can be found at https://knownagents.com/agents/shapbot |
| Sidetrade indexer bot | [Sidetrade](https://www.sidetrade.com) | Unclear at this time. | Extracts data for a variety of uses including training AI. | No information. | AI product training. | | Sidetrade indexer bot | [Sidetrade](https://www.sidetrade.com) | Unclear at this time. | Extracts data for a variety of uses including training AI. | No information. | AI product training. |
| Spider | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/spider |
| TavilyBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | TavilyBot is a web crawler by Tavily that indexes and extracts content from billions of pages, providing real-time search, extraction, and research data to ground AI agen… More info can be found at https://knownagents.com/agents/tavilybot |
| Terra Cotta | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Terra Cotta is Ceramic's web crawler that indexes public content to power their web-scale search API for AI and LLMs. More info can be found at https://knownagents.com/agents/terra-cotta |
| TerraCotta | [Ceramic AI](https://ceramic.ai/) | [Yes](https://github.com/CeramicTeam/CeramicTerracotta) | AI Data Providers | Unclear at this time. | TerraCotta is Ceramic's web crawler that indexes public content to power their web-scale search API for AI and LLMs. More info can be found at https://knownagents.com/agents/terracotta |
| Thinkbot | [Thinkbot](https://www.thinkbot.agency) | No | Insights on AI integration and automation. | Unclear at this time. | Collects data for analysis on AI usage and automation. |
| TikTokSpider | ByteDance | Unclear at this time. | LLM training. | Unclear at this time. | Downloads data to train LLMS, as per Bytespider. |
| Timpibot | [Timpi](https://timpi.io) | Unclear at this time. | Scrapes data for use in training LLMs. | No information. | Makes data available for training AI models. | | Timpibot | [Timpi](https://timpi.io) | Unclear at this time. | Scrapes data for use in training LLMs. | No information. | Makes data available for training AI models. |
| TongyiBot | Alibaba that fetches web content for the Tongyi Qianwen assistant and related Qwen-generated answers | Unclear at this time. | AI Assistants | Unclear at this time. | TongyiBot is a web crawler operated by Alibaba that fetches web content for the Tongyi Qianwen assistant and related Qwen-generated answers. More info can be found at https://knownagents.com/agents/tongyibot |
| Trae | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Trae is an AI-powered coding agent developed by ByteDance that can understand codebases, fetch web content, and generate code. More info can be found at https://knownagents.com/agents/trae |
| TwinAgent | Twin, a platform that creates automated workers to perform tasks by integrating with APIs and controlling web applications through browser automa… | Unclear at this time. | AI Agents | Unclear at this time. | TwinAgent is operated by Twin, a platform that creates automated workers to perform tasks by integrating with APIs and controlling web applications through browser automa… More info can be found at https://knownagents.com/agents/twinagent |
| UseAI | Unclear at this time. | Unclear at this time. | AI Assistants | Unclear at this time. | UseAI is a web crawler associated with Use AI, a platform that provides an AI workspace where users can chat with AI models, research the web, and perform various tasks. … More info can be found at https://knownagents.com/agents/useai |
| VelenPublicWebCrawler | [Velen Crawler](https://velen.io) | [Yes](https://velen.io) | Scrapes data for business data sets and machine learning models. | No information. | "Our goal with this crawler is to build business datasets and machine learning models to better understand the web." | | VelenPublicWebCrawler | [Velen Crawler](https://velen.io) | [Yes](https://velen.io) | Scrapes data for business data sets and machine learning models. | No information. | "Our goal with this crawler is to build business datasets and machine learning models to better understand the web." |
| Webzio-Extended | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://darkvisitors.com/agents/agents/webzio-extended | | WARDBot | WEBSPARK | Unclear at this time. | AI Data Scrapers | Unclear at this time. | WARDBot is an AI data scraper operated by WEBSPARK. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/wardbot |
| Webzio\-Extended | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://knownagents.com/agents/webzio-extended |
| webzio\-extended | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/webzio-extended |
| wpbot | [QuantumCloud](https://www.quantumcloud.com) | Unclear at this time; opt out provided via [Google Form](https://forms.gle/ajBaxygz9jSR8p8G9) | Live chat support and lead generation. | Unclear at this time. | wpbot is a used to support the functionality of the AI Chatbot for WordPress plugin. It supports the use of customer models, data collection and customer support. |
| WRTNBot | Unclear at this time. | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/wrtnbot |
| YaK | [Meltwater](https://www.meltwater.com/en/suite/consumer-intelligence) | Unclear at this time. | According to the [Meltwater Consumer Intelligence page](https://www.meltwater.com/en/suite/consumer-intelligence) 'By applying AI, data science, and market research expertise to a live feed of global data sources, we transform unstructured data into actionable insights allowing better decision-making'. | Unclear at this time. | Retrieves data used for Meltwater's AI enabled consumer intelligence suite |
| YandexAdditional | [Yandex](https://yandex.ru) | [Yes](https://yandex.ru/support/webmaster/en/search-appearance/fast.html?lang=en) | Scrapes/analyzes data for the YandexGPT LLM. | No information. | Retrieves data used for YandexGPT quick answers features. |
| YandexAdditionalBot | [Yandex](https://yandex.ru) | [Yes](https://yandex.ru/support/webmaster/en/search-appearance/fast.html?lang=en) | Scrapes/analyzes data for the YandexGPT LLM. | No information. | Retrieves data used for YandexGPT quick answers features. |
| YiyanBot | Baidu that fetches web content for the yiyan | Unclear at this time. | AI Assistants | Unclear at this time. | YiyanBot is a web crawler operated by Baidu that fetches web content for the yiyan.baidu.com assistant and related ERNIE-generated answers. More info can be found at https://knownagents.com/agents/yiyanbot |
| YouBot | [You](https://about.you.com/youchat/) | [Yes](https://about.you.com/youbot/) | Scrapes data for search engine and LLMs. | No information. | Retrieves data used for You.com web search engine and LLMs. | | YouBot | [You](https://about.you.com/youchat/) | [Yes](https://about.you.com/youbot/) | Scrapes data for search engine and LLMs. | No information. | Retrieves data used for You.com web search engine and LLMs. |
| ZanistaBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/zanistabot |