Compare commits

..

20 commits

Author SHA1 Message Date
Known Agents
987266f3c5 Update from knownagents.com 2026-09-26 02:41:07 +00:00
Glyn Normington
5631dee581
Merge pull request #289 from dunn/faq-entries
FAQ: add entries on labor and military use
2026-09-22 19:04:41 +01:00
alexandra catalina
1ecb571700 FAQ: add references subheading, convert to unordered list 2026-09-21 10:54:24 -07:00
alexandra catalina
57ed273dc7 FAQ: add entries on labor and military use 2026-09-19 11:13:23 -07:00
Glyn Normington
6c9c1a3327 Improve wording 2026-09-19 15:47:08 +01:00
Glyn Normington
fab855244f
Merge pull request #288 from ai-robots-txt/noAI
Do not permit AI contributions
2026-09-19 15:45:13 +01:00
Glyn Normington
e4c5d5e206 Do not permit AI contributions 2026-09-19 15:43:27 +01:00
Glyn Normington
425d1a6207
Merge pull request #287 from flober81/patch-1
Add AI Discovery Radar to related resources
2026-09-17 15:25:50 +01:00
Florian Berger
9f039df105
Shorten entry: drop dated figures, spell out the file types
Updated the description of the AI Discovery Radar dataset to clarify its purpose and measurement methodology.
2026-09-17 14:01:37 +02:00
Florian Berger
8136c08d6a
Add AI Discovery Radar to related resources 2026-09-17 10:47:31 +02:00
Glyn Normington
0e111dcc24
Merge pull request #285 from SamHartleyFixes/match-agent-names-case-insensitively
Resolve agent names case-insensitively when ingesting knownagents.com
2026-09-08 06:20:04 +01:00
ai.robots.txt
197f156dd9 Merge pull request #286 from SamHartleyFixes/case-insensitive-matching
Match generated nginx/lighttpd/Caddy configs case-insensitively
2026-09-08 05:19:31 +00:00
Glyn Normington
ebc94b445f
Merge pull request #286 from SamHartleyFixes/case-insensitive-matching
Match generated nginx/lighttpd/Caddy configs case-insensitively
2026-09-08 06:19:20 +01:00
Sam Hartley
86a2e1fc67 Regenerate Caddyfile test fixture for case-insensitive match 2026-09-07 20:08:07 -07:00
Sam Hartley
3adc0cb1db Regenerate lighttpd test fixture for case-insensitive match 2026-09-07 20:08:06 -07:00
Sam Hartley
aa623355c1 Regenerate nginx test fixture for case-insensitive match 2026-09-07 20:08:06 -07:00
Sam Hartley
447bb15a22 Make the nginx/lighttpd/Caddy configs match case-insensitively 2026-09-07 20:08:05 -07:00
SamHartleyFixes
e028137f46 Resolve agent names case-insensitively when ingesting knownagents.com
robots.txt user-agent matching is case-insensitive, so two entries whose
names differ only in case are the same crawler. updated_robots_json keyed
off the scraped name, so when knownagents.com changed the capitalisation
of a name the ingest added a second entry instead of updating the first.
consolidate() also looks the name up by exact key, so the curated operator
and respect values were left behind on the old key rather than carried over.

robots.json currently carries three such pairs, and each one emits a
duplicate User-agent line in every generated file.

This resolves an incoming name against the keys already present before
using it, so an existing entry is updated in place. No generated file
changes as a result.
2026-09-07 18:08:23 -07:00
Glyn Normington
db83172548
Merge pull request #283 from SamHartleyFixes/ai-crawler-census
Add AI Crawler Census to Related
2026-09-07 18:08:07 +01:00
Sam Hartley
ee9805ef9f Add AI Crawler Census to Related 2026-09-07 08:09:25 -07:00
15 changed files with 132 additions and 21 deletions

View file

@ -1,3 +1,3 @@
RewriteEngine On RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b) [NC] RewriteCond %{HTTP_USER_AGENT} (\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b) [NC]
RewriteRule !^/?robots\.txt$ - [F] RewriteRule !^/?robots\.txt$ - [F]

View file

@ -1,3 +1,3 @@
@aibots { @aibots {
header_regexp User-Agent "\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b" header_regexp User-Agent "(?i)\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b"
} }

34
FAQ.md
View file

@ -2,19 +2,35 @@
## Why should we block these crawlers? ## Why should we block these crawlers?
They're extractive, confer no benefit to the creators of data they're ingesting and also have wide-ranging negative externalities: particularly copyright abuse and environmental impact. They're extractive, confer no benefit to the creators of data they're ingesting
and also have wide-ranging negative externalities, from copyright abuse and
environmental impacts to the exploitation of labor and use in war.
**[How Tech Giants Cut Corners to Harvest Data for A.I.](https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html?unlocked_article_code=1.ik0.Ofja.L21c1wyW-0xj&ugrp=m)** ### References
> OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems.
**[How AI copyright lawsuits could make the whole industry go extinct](https://www.theverge.com/24062159/ai-copyright-fair-use-lawsuits-new-york-times-openai-chatgpt-decoder-podcast)** - [How Tech Giants Cut Corners to Harvest Data for A.I.](https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html?unlocked_article_code=1.ik0.Ofja.L21c1wyW-0xj&ugrp=m)
> The New York Times' lawsuit against OpenAI is part of a broader, industry-shaking copyright challenge that could define the future of AI.
**[Reconciling the contrasting narratives on the environmental impact of large language models](https://www.nature.com/articles/s41598-024-76682-6)** > OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems.
> Studies have shown that the training of just one LLM can consume as much energy as five cars do across their lifetimes. The water footprint of AI is also substantial; for example, recent work has highlighted that water consumption associated with AI models involves data centers using millions of gallons of water per day for cooling. Additionally, the energy consumption and carbon emissions of AI are projected to grow quickly in the coming years [...].
**[Scientists Predict AI to Generate Millions of Tons of E-Waste](https://www.sciencealert.com/scientists-predict-ai-to-generate-millions-of-tons-of-e-waste)** - [How AI copyright lawsuits could make the whole industry go extinct](https://www.theverge.com/24062159/ai-copyright-fair-use-lawsuits-new-york-times-openai-chatgpt-decoder-podcast)
> we could end up with between 1.2 million and 5 million metric tons of additional electronic waste by the end of this decade [the 2020's].
> The New York Times' lawsuit against OpenAI is part of a broader, industry-shaking copyright challenge that could define the future of AI.
- [Reconciling the contrasting narratives on the environmental impact of large language models](https://www.nature.com/articles/s41598-024-76682-6)
> Studies have shown that the training of just one LLM can consume as much energy as five cars do across their lifetimes. The water footprint of AI is also substantial; for example, recent work has highlighted that water consumption associated with AI models involves data centers using millions of gallons of water per day for cooling. Additionally, the energy consumption and carbon emissions of AI are projected to grow quickly in the coming years [...].
- [Scientists Predict AI to Generate Millions of Tons of E-Waste](https://www.sciencealert.com/scientists-predict-ai-to-generate-millions-of-tons-of-e-waste)
> we could end up with between 1.2 million and 5 million metric tons of additional electronic waste by the end of this decade [the 2020's].
- [Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic](https://time.com/6247678/openai-chatgpt-kenya-workers/)
> AI often relies on hidden human labor in the Global South that can often be damaging and exploitative. These invisible workers remain on the margins even as their work contributes to billion-dollar industries.
- [US Military Using Claude to Select Targets in Iran Strikes](https://futurism.com/artificial-intelligence/claude-anthropic-military-iran)
> Anthropic’s large language model, Claude, is the key “AI tool” used by US Central Command in the Middle East. Its tasks include assessing intelligence, simulated war games, and even identifying military targets — in short, helping military leaders plan attacks that have already claimed hundreds of lives.
## How do we know AI companies/bots respect `robots.txt`? ## How do we know AI companies/bots respect `robots.txt`?

View file

@ -56,8 +56,14 @@ file on-the-fly.
- [KI-Zugangsindex](https://peppe1337.github.io/ki-zugangsindex/): open dataset on how widely this kind of blocking is actually deployed in the German (`.de`) web, measured on a fixed panel of 600 domains so the same sites can be re-checked over time. - [KI-Zugangsindex](https://peppe1337.github.io/ki-zugangsindex/): open dataset on how widely this kind of blocking is actually deployed in the German (`.de`) web, measured on a fixed panel of 600 domains so the same sites can be re-checked over time.
- [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0.
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0.
## Contributing ## Contributing
Please note that AI-generated contributions are not permitted.
A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`. A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`.
You can run the tests by [installing](https://www.python.org/about/gettingstarted/) Python 3, installing the dependencies: You can run the tests by [installing](https://www.python.org/about/gettingstarted/) Python 3, installing the dependencies:

View file

@ -34,6 +34,25 @@ default_values = {
} }
default_value = "Unclear at this time." default_value = "Unclear at this time."
def existing_key(existing_content, name: str) -> str:
"""Return the key robots.json already uses for this agent, ignoring case.
robots.txt user-agent matching is case-insensitive, so two entries whose
names differ only in case are the same crawler. knownagents.com has changed
the capitalisation of a name before now, and keying off the scraped name
added a second entry rather than updating the first, which both duplicated
the User-agent line in every generated file and left the curated operator
and respect values behind on the old key.
"""
if name in existing_content:
return name
lowered = name.lower()
for key in existing_content:
if key.lower() == lowered:
return key
return name
def consolidate(existing_content, name: str, field: str, value: str) -> str: def consolidate(existing_content, name: str, field: str, value: str) -> str:
# New entry # New entry
if name not in existing_content: if name not in existing_content:
@ -77,6 +96,7 @@ def updated_robots_json(soup):
for agent in section.find_all("a", href=True): for agent in section.find_all("a", href=True):
name = agent.find("div", {"class": "agent-name"}).get_text().strip() name = agent.find("div", {"class": "agent-name"}).get_text().strip()
name = clean_robot_name(name) name = clean_robot_name(name)
name = existing_key(existing_content, name)
desc_tag = agent.find("div", {"class": "description"}) desc_tag = agent.find("div", {"class": "description"})
if desc_tag is not None: if desc_tag is not None:
@ -201,7 +221,9 @@ def json_to_htaccess(robot_json):
def json_to_nginx(robot_json): def json_to_nginx(robot_json):
# Creates an Nginx config file. This config snippet can be included in # Creates an Nginx config file. This config snippet can be included in
# nginx server{} blocks to block AI bots. # nginx server{} blocks to block AI bots.
config = f"set $block 0;\n\nif ($http_user_agent ~ {list_to_pcre(robot_json)!r}) {{\n set $block 1;\n}}\n\nif ($request_uri = '/robots.txt') {{\n set $block 0;\n}}\n\nif ($block) {{\n return 403;\n}}" # Exact User-agent matching is case-insensitive (RFC 9309 2.2.1), so this uses
# nginx's case-insensitive "~*" operator rather than "~".
config = f"set $block 0;\n\nif ($http_user_agent ~* {list_to_pcre(robot_json)!r}) {{\n set $block 1;\n}}\n\nif ($request_uri = '/robots.txt') {{\n set $block 0;\n}}\n\nif ($block) {{\n return 403;\n}}"
return config return config
@ -210,7 +232,9 @@ def json_to_lighttpd(robot_json):
# Lighttpd configuration global or in $HTTP conditionals to block AI bots. # Lighttpd configuration global or in $HTTP conditionals to block AI bots.
# single quotes (as returned by repr) are not valid string delimeters, so we # single quotes (as returned by repr) are not valid string delimeters, so we
# must manually quote it end ensure no unescaped quotes are inside. # must manually quote it end ensure no unescaped quotes are inside.
escaped_quotes = list_to_pcre(robot_json).replace('"', '\\"') # Lighttpd's "=~" is case-sensitive by default; the inline (?i) flag makes it
# match the case-insensitive exact matching robots.txt itself requires.
escaped_quotes = ("(?i)" + list_to_pcre(robot_json)).replace('"', '\\"')
config = f'$HTTP["url"] != "/robots.txt" {{ $HTTP["user-agent"] =~ "{escaped_quotes}" {{ url.access-deny = ( "" ) }} }}' config = f'$HTTP["url"] != "/robots.txt" {{ $HTTP["user-agent"] =~ "{escaped_quotes}" {{ url.access-deny = ( "" ) }} }}'
return config return config
@ -218,7 +242,8 @@ def json_to_lighttpd(robot_json):
def json_to_caddy(robot_json): def json_to_caddy(robot_json):
# single quotes (as returned by repr) are not valid string delimeters, so we # single quotes (as returned by repr) are not valid string delimeters, so we
# must manually quote it end ensure no unescaped quotes are inside. # must manually quote it end ensure no unescaped quotes are inside.
escaped_quotes = list_to_pcre(robot_json).replace('"', '\\"') # Same case-insensitivity note as lighttpd; Caddy's RE2 engine also honours (?i).
escaped_quotes = ("(?i)" + list_to_pcre(robot_json)).replace('"', '\\"')
caddyfile = "@aibots {\n " caddyfile = "@aibots {\n "
caddyfile += f' header_regexp User-Agent "{escaped_quotes}"' caddyfile += f' header_regexp User-Agent "{escaped_quotes}"'
caddyfile += "\n}" caddyfile += "\n}"

View file

@ -1,3 +1,3 @@
@aibots { @aibots {
header_regexp User-Agent "\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b" header_regexp User-Agent "(?i)\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b"
} }

View file

@ -1 +1 @@
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b" { url.access-deny = ( "" ) } } $HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "(?i)\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b" { url.access-deny = ( "" ) } }

View file

@ -1,6 +1,6 @@
set $block 0; set $block 0;
if ($http_user_agent ~ '\\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\\ indexer\\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\\.with\\.dots|star\\*\\*\\*crawler|Is\\ this\\ a\\ crawler\\?|a\\[mazing\\]\\{42\\}\\(robot\\)|2\\^32\\$|curl\\|sudo\\ bash)\\b') { if ($http_user_agent ~* '\\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\\ indexer\\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\\.with\\.dots|star\\*\\*\\*crawler|Is\\ this\\ a\\ crawler\\?|a\\[mazing\\]\\{42\\}\\(robot\\)|2\\^32\\$|curl\\|sudo\\ bash)\\b') {
set $block 1; set $block 1;
} }

View file

@ -7,6 +7,7 @@ import unittest
from robots import ( from robots import (
consolidate, consolidate,
existing_key,
default_value, default_value,
default_values, default_values,
json_to_caddy, json_to_caddy,
@ -203,6 +204,19 @@ class TestConsolidate(unittest.TestCase, RobotsUnittestExtensions):
self.assertEqual("Rosie is the robot maid from The Jetsons, an American animated sitcom", self.assertEqual("Rosie is the robot maid from The Jetsons, an American animated sitcom",
consolidate(existing, "rosie", "description", "Rosie is the robot maid from The Jetsons, an American animated sitcom")) consolidate(existing, "rosie", "description", "Rosie is the robot maid from The Jetsons, an American animated sitcom"))
class TestExistingKey(unittest.TestCase):
def test_exact_match_wins(self):
existing = {"Rosie": {}, "rosie": {}}
self.assertEqual("rosie", existing_key(existing, "rosie"))
def test_matches_ignoring_case(self):
existing = {"Rosie": {"operator": "George Jetson"}}
self.assertEqual("Rosie", existing_key(existing, "rosie"))
def test_unknown_name_is_returned_unchanged(self):
self.assertEqual("rosie", existing_key({}, "rosie"))
if __name__ == "__main__": if __name__ == "__main__":
import os import os
os.chdir(os.path.dirname(__file__)) os.chdir(os.path.dirname(__file__))

View file

@ -1,4 +1,5 @@
AddSearchBot AddSearchBot
AgentDataBot
AgentTimes AgentTimes
AI2Bot AI2Bot
AI2Bot-DeepResearchEval AI2Bot-DeepResearchEval
@ -24,6 +25,7 @@ Awario
AzureAI-SearchBot AzureAI-SearchBot
bedrockbot bedrockbot
bigsur.ai bigsur.ai
BixelBot
Bravebot Bravebot
Brightbot Brightbot
Brightbot 1.0 Brightbot 1.0
@ -40,6 +42,7 @@ Claude-User
Claude-Web Claude-Web
ClaudeBot ClaudeBot
Cloudflare-AutoRAG Cloudflare-AutoRAG
CloudflareBrowserRenderingCrawler
CloudVertexBot CloudVertexBot
Code Code
cohere-ai cohere-ai
@ -91,6 +94,7 @@ img2dataset
ISSCyberRiskCrawler ISSCyberRiskCrawler
kagi-fetcher kagi-fetcher
Kangaroo Bot Kangaroo Bot
Kimi-Agent
Kimi-SearchBot Kimi-SearchBot
Kimi-User Kimi-User
KimiBot KimiBot
@ -136,6 +140,7 @@ PetalBot
PhindBot PhindBot
Poggio-Citations Poggio-Citations
Poseidon Research Crawler Poseidon Research Crawler
qodercli
QualifiedBot QualifiedBot
Querit-SearchBot Querit-SearchBot
QueritBot QueritBot

View file

@ -1 +1 @@
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b" { url.access-deny = ( "" ) } } $HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "(?i)\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b" { url.access-deny = ( "" ) } }

View file

@ -1,6 +1,6 @@
set $block 0; set $block 0;
if ($http_user_agent ~ '\\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\\.ai|Bravebot|Brightbot|Brightbot\\ 1\\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\\ Imprint\\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\\ Research\\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\\ indexer\\ bot|Spider|TavilyBot|Terra\\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\\b') { if ($http_user_agent ~* '\\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\\.ai|BixelBot|Bravebot|Brightbot|Brightbot\\ 1\\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\\ Imprint\\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\\ Research\\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\\ indexer\\ bot|Spider|TavilyBot|Terra\\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\\b') {
set $block 1; set $block 1;
} }

View file

@ -6,6 +6,13 @@
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot" "description": "AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot"
}, },
"AgentDataBot": {
"operator": "Unclear at this time.",
"respect": "Unclear at this time.",
"function": "AI Data Providers",
"frequency": "Unclear at this time.",
"description": "AgentDataBot visits public web pages to identify technology stacks, company details, and team members. The data builds AgentData's company-intelligence database for AI ag\u2026 More info can be found at https://knownagents.com/agents/agentdatabot"
},
"AgentTimes": { "AgentTimes": {
"operator": "[The Agent Times](https://theagenttimes.com/about)", "operator": "[The Agent Times](https://theagenttimes.com/about)",
"respect": "Unclear at this time.", "respect": "Unclear at this time.",
@ -181,6 +188,13 @@
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi\u2026 More info can be found at https://knownagents.com/agents/bigsur-ai" "description": "bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi\u2026 More info can be found at https://knownagents.com/agents/bigsur-ai"
}, },
"BixelBot": {
"operator": "Unclear at this time.",
"respect": "Unclear at this time.",
"function": "AI Data Providers",
"frequency": "Unclear at this time.",
"description": "BixelBot collects public company and market signals for Bixel's structured company-data platform, which supplies verified business information to AI agents and developers\u2026 More info can be found at https://knownagents.com/agents/bixelbot"
},
"Bravebot": { "Bravebot": {
"operator": "https://safe.search.brave.com/help/brave-search-crawler", "operator": "https://safe.search.brave.com/help/brave-search-crawler",
"respect": "Yes", "respect": "Yes",
@ -293,6 +307,13 @@
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "AutoRAG is an all-in-one AI search solution." "description": "AutoRAG is an all-in-one AI search solution."
}, },
"CloudflareBrowserRenderingCrawler": {
"operator": "Cloudflare that returns rendered website content for research, monitoring, and AI data workflows",
"respect": "Unclear at this time.",
"function": "AI Data Providers",
"frequency": "Unclear at this time.",
"description": "CloudflareBrowserRenderingCrawler is a web crawler operated by Cloudflare that returns rendered website content for research, monitoring, and AI data workflows. More info can be found at https://knownagents.com/agents/cloudflarebrowserrenderingcrawler"
},
"CloudVertexBot": { "CloudVertexBot": {
"operator": "Unclear at this time.", "operator": "Unclear at this time.",
"respect": "Unclear at this time.", "respect": "Unclear at this time.",
@ -651,6 +672,13 @@
"frequency": "Unclear at this time.", "frequency": "Unclear at this time.",
"description": "Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot" "description": "Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot"
}, },
"Kimi-Agent": {
"operator": "Moonshot AI that browses websites while carrying out user-directed research and other tasks",
"respect": "Unclear at this time.",
"function": "AI Agents",
"frequency": "Unclear at this time.",
"description": "Kimi-Agent is a browser-enabled AI agent operated by Moonshot AI that browses websites while carrying out user-directed research and other tasks. More info can be found at https://knownagents.com/agents/kimi-agent"
},
"Kimi-SearchBot": { "Kimi-SearchBot": {
"operator": "[Moonshot AI](https://www.moonshot.ai)", "operator": "[Moonshot AI](https://www.moonshot.ai)",
"respect": "[Yes](https://www.kimi.ai/policies/kimi-crawlers)", "respect": "[Yes](https://www.kimi.ai/policies/kimi-crawlers)",
@ -966,6 +994,13 @@
"function": "AI research crawler", "function": "AI research crawler",
"respect": "Unclear at this time." "respect": "Unclear at this time."
}, },
"qodercli": {
"operator": "Qoder that helps developers build, debug, and modify software from the terminal",
"respect": "Unclear at this time.",
"function": "AI Coding Agents",
"frequency": "Unclear at this time.",
"description": "qodercli is a command-line AI coding agent operated by Qoder that helps developers build, debug, and modify software from the terminal. More info can be found at https://knownagents.com/agents/qodercli"
},
"QualifiedBot": { "QualifiedBot": {
"operator": "[Qualified](https://www.qualified.com)", "operator": "[Qualified](https://www.qualified.com)",
"respect": "Unclear at this time.", "respect": "Unclear at this time.",

View file

@ -1,4 +1,5 @@
User-agent: AddSearchBot User-agent: AddSearchBot
User-agent: AgentDataBot
User-agent: AgentTimes User-agent: AgentTimes
User-agent: AI2Bot User-agent: AI2Bot
User-agent: AI2Bot-DeepResearchEval User-agent: AI2Bot-DeepResearchEval
@ -24,6 +25,7 @@ User-agent: Awario
User-agent: AzureAI-SearchBot User-agent: AzureAI-SearchBot
User-agent: bedrockbot User-agent: bedrockbot
User-agent: bigsur.ai User-agent: bigsur.ai
User-agent: BixelBot
User-agent: Bravebot User-agent: Bravebot
User-agent: Brightbot User-agent: Brightbot
User-agent: Brightbot 1.0 User-agent: Brightbot 1.0
@ -40,6 +42,7 @@ User-agent: Claude-User
User-agent: Claude-Web User-agent: Claude-Web
User-agent: ClaudeBot User-agent: ClaudeBot
User-agent: Cloudflare-AutoRAG User-agent: Cloudflare-AutoRAG
User-agent: CloudflareBrowserRenderingCrawler
User-agent: CloudVertexBot User-agent: CloudVertexBot
User-agent: Code User-agent: Code
User-agent: cohere-ai User-agent: cohere-ai
@ -91,6 +94,7 @@ User-agent: img2dataset
User-agent: ISSCyberRiskCrawler User-agent: ISSCyberRiskCrawler
User-agent: kagi-fetcher User-agent: kagi-fetcher
User-agent: Kangaroo Bot User-agent: Kangaroo Bot
User-agent: Kimi-Agent
User-agent: Kimi-SearchBot User-agent: Kimi-SearchBot
User-agent: Kimi-User User-agent: Kimi-User
User-agent: KimiBot User-agent: KimiBot
@ -136,6 +140,7 @@ User-agent: PetalBot
User-agent: PhindBot User-agent: PhindBot
User-agent: Poggio-Citations User-agent: Poggio-Citations
User-agent: Poseidon Research Crawler User-agent: Poseidon Research Crawler
User-agent: qodercli
User-agent: QualifiedBot User-agent: QualifiedBot
User-agent: Querit-SearchBot User-agent: Querit-SearchBot
User-agent: QueritBot User-agent: QueritBot

View file

@ -1,6 +1,7 @@
| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description | | Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |
|------|----------|-----------------------|----------|------------------|-------------| |------|----------|-----------------------|----------|------------------|-------------|
| AddSearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot | | AddSearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot |
| AgentDataBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | AgentDataBot visits public web pages to identify technology stacks, company details, and team members. The data builds AgentData's company-intelligence database for AI ag… More info can be found at https://knownagents.com/agents/agentdatabot |
| AgentTimes | [The Agent Times](https://theagenttimes.com/about) | Unclear at this time. | Data Scraper from RSS Feeds. | Requests RSS feed every 5-6 minutes. | Scrapes data for AI news aggregation and republishing. | | AgentTimes | [The Agent Times](https://theagenttimes.com/about) | Unclear at this time. | Data Scraper from RSS Feeds. | Requests RSS feed every 5-6 minutes. | Scrapes data for AI news aggregation and republishing. |
| AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. | | AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
| AI2Bot\-DeepResearchEval | Ai2, a non-profit AI research institute | Unclear at this time. | AI Assistants | Unclear at this time. | Ai2Bot-DeepResearchEval is operated by Ai2, a non-profit AI research institute. It's used to collect and scan resources used in deep research queries performed by Ai2's o… More info can be found at https://knownagents.com/agents/ai2bot-deepresearcheval | | AI2Bot\-DeepResearchEval | Ai2, a non-profit AI research institute | Unclear at this time. | AI Assistants | Unclear at this time. | Ai2Bot-DeepResearchEval is operated by Ai2, a non-profit AI research institute. It's used to collect and scan resources used in deep research queries performed by Ai2's o… More info can be found at https://knownagents.com/agents/ai2bot-deepresearcheval |
@ -26,6 +27,7 @@
| AzureAI\-SearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/azureai-searchbot | | AzureAI\-SearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/azureai-searchbot |
| bedrockbot | [Amazon](https://amazon.com) | [Yes](https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html#configuration-webcrawl-connector) | Data scraping for custom AI applications. | Unclear at this time. | Connects to and crawls URLs that have been selected for use in a user's AWS bedrock application. | | bedrockbot | [Amazon](https://amazon.com) | [Yes](https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html#configuration-webcrawl-connector) | Data scraping for custom AI applications. | Unclear at this time. | Connects to and crawls URLs that have been selected for use in a user's AWS bedrock application. |
| bigsur\.ai | Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… | Unclear at this time. | AI Assistants | Unclear at this time. | bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… More info can be found at https://knownagents.com/agents/bigsur-ai | | bigsur\.ai | Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… | Unclear at this time. | AI Assistants | Unclear at this time. | bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… More info can be found at https://knownagents.com/agents/bigsur-ai |
| BixelBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | BixelBot collects public company and market signals for Bixel's structured company-data platform, which supplies verified business information to AI agents and developers… More info can be found at https://knownagents.com/agents/bixelbot |
| Bravebot | https://safe.search.brave.com/help/brave-search-crawler | Yes | AI Data Providers | Unclear at this time. | Bravebot is a web crawler by Brave that indexes pages for Brave Search, providing search data and AI-optimized context to power chatbots, agents, and RAG pipelines. More info can be found at https://knownagents.com/agents/bravebot | | Bravebot | https://safe.search.brave.com/help/brave-search-crawler | Yes | AI Data Providers | Unclear at this time. | Bravebot is a web crawler by Brave that indexes pages for Brave Search, providing search data and AI-optimized context to power chatbots, agents, and RAG pipelines. More info can be found at https://knownagents.com/agents/bravebot |
| Brightbot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Brightbot is a web data collection crawler by Bright Data that extracts and structures public website content at scale, providing AI-ready data for model training, RAG pi… More info can be found at https://knownagents.com/agents/brightbot | | Brightbot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Brightbot is a web data collection crawler by Bright Data that extracts and structures public website content at scale, providing AI-ready data for model training, RAG pi… More info can be found at https://knownagents.com/agents/brightbot |
| Brightbot 1\.0 | https://brightdata.com/brightbot | Unclear at this time. | LLM/AI training. | At least one per minute. | Scrapes data to train LLMs and AI products focused on website customer support, [uses residential IPs and legit-looking user-agents to disguise itself](https://ksol.io/en/blog/posts/brightbot-not-that-bright/). | | Brightbot 1\.0 | https://brightdata.com/brightbot | Unclear at this time. | LLM/AI training. | At least one per minute. | Scrapes data to train LLMs and AI products focused on website customer support, [uses residential IPs and legit-looking user-agents to disguise itself](https://ksol.io/en/blog/posts/brightbot-not-that-bright/). |
@ -42,6 +44,7 @@
| Claude\-Web | Anthropic | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Claude-Web is an AI-related agent operated by Anthropic. It's currently unclear exactly what it's used for, since there's no official documentation. If you can provide more detail, please contact us. More info can be found at https://knownagents.com/agents/claude-web | | Claude\-Web | Anthropic | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Claude-Web is an AI-related agent operated by Anthropic. It's currently unclear exactly what it's used for, since there's no official documentation. If you can provide more detail, please contact us. More info can be found at https://knownagents.com/agents/claude-web |
| ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. | | ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
| Cloudflare\-AutoRAG | [Cloudflare](https://developers.cloudflare.com/autorag) | Yes | Collects data for AI search | Unclear at this time. | AutoRAG is an all-in-one AI search solution. | | Cloudflare\-AutoRAG | [Cloudflare](https://developers.cloudflare.com/autorag) | Yes | Collects data for AI search | Unclear at this time. | AutoRAG is an all-in-one AI search solution. |
| CloudflareBrowserRenderingCrawler | Cloudflare that returns rendered website content for research, monitoring, and AI data workflows | Unclear at this time. | AI Data Providers | Unclear at this time. | CloudflareBrowserRenderingCrawler is a web crawler operated by Cloudflare that returns rendered website content for research, monitoring, and AI data workflows. More info can be found at https://knownagents.com/agents/cloudflarebrowserrenderingcrawler |
| CloudVertexBot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform. More info can be found at https://knownagents.com/agents/cloudvertexbot | | CloudVertexBot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform. More info can be found at https://knownagents.com/agents/cloudvertexbot |
| Code | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Code (GitHub Copilot) is an AI coding agent that can autonomously plan, build, and execute development tasks, functioning as a collaborative AI pair programmer. More info can be found at https://knownagents.com/agents/code | | Code | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Code (GitHub Copilot) is an AI coding agent that can autonomously plan, build, and execute development tasks, functioning as a collaborative AI pair programmer. More info can be found at https://knownagents.com/agents/code |
| cohere\-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. | | cohere\-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. |
@ -93,6 +96,7 @@
| ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. | | ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. |
| kagi\-fetcher | Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… | Unclear at this time. | AI Assistants | Unclear at this time. | kagi-fetcher is an AI Assistant operated by Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… More info can be found at https://knownagents.com/agents/kagi-fetcher | | kagi\-fetcher | Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… | Unclear at this time. | AI Assistants | Unclear at this time. | kagi-fetcher is an AI Assistant operated by Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… More info can be found at https://knownagents.com/agents/kagi-fetcher |
| Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot | | Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot |
| Kimi\-Agent | Moonshot AI that browses websites while carrying out user-directed research and other tasks | Unclear at this time. | AI Agents | Unclear at this time. | Kimi-Agent is a browser-enabled AI agent operated by Moonshot AI that browses websites while carrying out user-directed research and other tasks. More info can be found at https://knownagents.com/agents/kimi-agent |
| Kimi\-SearchBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Search Crawlers | No information provided. | Kimi-SearchBot powers Kimi's search features: it analyzes pages for relevance and builds the search index. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers | | Kimi\-SearchBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Search Crawlers | No information provided. | Kimi-SearchBot powers Kimi's search features: it analyzes pages for relevance and builds the search index. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
| Kimi\-User | Moonshot AI that fetches web content on behalf of users interacting with Kimi | Unclear at this time. | AI Assistants | Unclear at this time. | Kimi-User is a web crawler operated by Moonshot AI that fetches web content on behalf of users interacting with Kimi. When a user asks Kimi to summarize an article or ans… More info can be found at https://knownagents.com/agents/kimi-user | | Kimi\-User | Moonshot AI that fetches web content on behalf of users interacting with Kimi | Unclear at this time. | AI Assistants | Unclear at this time. | Kimi-User is a web crawler operated by Moonshot AI that fetches web content on behalf of users interacting with Kimi. When a user asks Kimi to summarize an article or ans… More info can be found at https://knownagents.com/agents/kimi-user |
| KimiBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Data Scrapers | No information provided. | KimiBot crawls content potentially used to train Kimi's foundation models. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers | | KimiBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Data Scrapers | No information provided. | KimiBot crawls content potentially used to train Kimi's foundation models. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
@ -138,6 +142,7 @@
| PhindBot | [phind](https://www.phind.com/) | Unclear at this time. | AI Assistants | No explicit frequency provided. | Phind is an AI-powered answer engine designed for developers, offering technical answers and code examples. It uses real-time web search and specialized AI models to prov… More info can be found at https://knownagents.com/agents/phindbot | | PhindBot | [phind](https://www.phind.com/) | Unclear at this time. | AI Assistants | No explicit frequency provided. | Phind is an AI-powered answer engine designed for developers, offering technical answers and code examples. It uses real-time web search and specialized AI models to prov… More info can be found at https://knownagents.com/agents/phindbot |
| Poggio\-Citations | Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… | Unclear at this time. | AI Assistants | Unclear at this time. | Poggio-Citations is a web crawler operated by Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… More info can be found at https://knownagents.com/agents/poggio-citations | | Poggio\-Citations | Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… | Unclear at this time. | AI Assistants | Unclear at this time. | Poggio-Citations is a web crawler operated by Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… More info can be found at https://knownagents.com/agents/poggio-citations |
| Poseidon Research Crawler | [Poseidon Research](https://www.poseidonresearch.com) | Unclear at this time. | AI research crawler | No explicit frequency provided. | Lab focused on scaling the interpretability research necessary to make better AI systems possible. | | Poseidon Research Crawler | [Poseidon Research](https://www.poseidonresearch.com) | Unclear at this time. | AI research crawler | No explicit frequency provided. | Lab focused on scaling the interpretability research necessary to make better AI systems possible. |
| qodercli | Qoder that helps developers build, debug, and modify software from the terminal | Unclear at this time. | AI Coding Agents | Unclear at this time. | qodercli is a command-line AI coding agent operated by Qoder that helps developers build, debug, and modify software from the terminal. More info can be found at https://knownagents.com/agents/qodercli |
| QualifiedBot | [Qualified](https://www.qualified.com) | Unclear at this time. | AI Assistants | No explicit frequency provided. | QualifiedBot is Qualified's web crawler that analyzes customer websites to provide contextual information for their AI-powered chatbots and conversational marketing platf… More info can be found at https://knownagents.com/agents/qualifiedbot | | QualifiedBot | [Qualified](https://www.qualified.com) | Unclear at this time. | AI Assistants | No explicit frequency provided. | QualifiedBot is Qualified's web crawler that analyzes customer websites to provide contextual information for their AI-powered chatbots and conversational marketing platf… More info can be found at https://knownagents.com/agents/qualifiedbot |
| Querit\-SearchBot | Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… | Unclear at this time. | AI Data Providers | Unclear at this time. | Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… More info can be found at https://knownagents.com/agents/querit-searchbot | | Querit\-SearchBot | Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… | Unclear at this time. | AI Data Providers | Unclear at this time. | Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… More info can be found at https://knownagents.com/agents/querit-searchbot |
| QueritBot | Querit, a company providing a search API for large language model integration | Unclear at this time. | AI Data Providers | Unclear at this time. | QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time … More info can be found at https://knownagents.com/agents/queritbot | | QueritBot | Querit, a company providing a search API for large language model integration | Unclear at this time. | AI Data Providers | Unclear at this time. | QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time … More info can be found at https://knownagents.com/agents/queritbot |