mirror of
https://github.com/ai-robots-txt/ai.robots.txt.git
synced 2026-09-29 13:24:15 +02:00
Compare commits
20 commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
987266f3c5 | ||
|
|
5631dee581 |
||
|
|
1ecb571700 | ||
|
|
57ed273dc7 | ||
|
|
6c9c1a3327 | ||
|
|
fab855244f |
||
|
|
e4c5d5e206 | ||
|
|
425d1a6207 |
||
|
|
9f039df105 |
||
|
|
8136c08d6a |
||
|
|
0e111dcc24 |
||
|
|
197f156dd9 | ||
|
|
ebc94b445f |
||
|
|
86a2e1fc67 | ||
|
|
3adc0cb1db | ||
|
|
aa623355c1 | ||
|
|
447bb15a22 | ||
|
|
e028137f46 | ||
|
|
db83172548 |
||
|
|
ee9805ef9f |
15 changed files with 132 additions and 21 deletions
|
|
@ -1,3 +1,3 @@
|
||||||
RewriteEngine On
|
RewriteEngine On
|
||||||
RewriteCond %{HTTP_USER_AGENT} (\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b) [NC]
|
RewriteCond %{HTTP_USER_AGENT} (\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b) [NC]
|
||||||
RewriteRule !^/?robots\.txt$ - [F]
|
RewriteRule !^/?robots\.txt$ - [F]
|
||||||
|
|
|
||||||
|
|
@ -1,3 +1,3 @@
|
||||||
@aibots {
|
@aibots {
|
||||||
header_regexp User-Agent "\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b"
|
header_regexp User-Agent "(?i)\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b"
|
||||||
}
|
}
|
||||||
34
FAQ.md
34
FAQ.md
|
|
@ -2,19 +2,35 @@
|
||||||
|
|
||||||
## Why should we block these crawlers?
|
## Why should we block these crawlers?
|
||||||
|
|
||||||
They're extractive, confer no benefit to the creators of data they're ingesting and also have wide-ranging negative externalities: particularly copyright abuse and environmental impact.
|
They're extractive, confer no benefit to the creators of data they're ingesting
|
||||||
|
and also have wide-ranging negative externalities, from copyright abuse and
|
||||||
|
environmental impacts to the exploitation of labor and use in war.
|
||||||
|
|
||||||
**[How Tech Giants Cut Corners to Harvest Data for A.I.](https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html?unlocked_article_code=1.ik0.Ofja.L21c1wyW-0xj&ugrp=m)**
|
### References
|
||||||
> OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems.
|
|
||||||
|
|
||||||
**[How AI copyright lawsuits could make the whole industry go extinct](https://www.theverge.com/24062159/ai-copyright-fair-use-lawsuits-new-york-times-openai-chatgpt-decoder-podcast)**
|
- [How Tech Giants Cut Corners to Harvest Data for A.I.](https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html?unlocked_article_code=1.ik0.Ofja.L21c1wyW-0xj&ugrp=m)
|
||||||
> The New York Times' lawsuit against OpenAI is part of a broader, industry-shaking copyright challenge that could define the future of AI.
|
|
||||||
|
|
||||||
**[Reconciling the contrasting narratives on the environmental impact of large language models](https://www.nature.com/articles/s41598-024-76682-6)**
|
> OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems.
|
||||||
> Studies have shown that the training of just one LLM can consume as much energy as five cars do across their lifetimes. The water footprint of AI is also substantial; for example, recent work has highlighted that water consumption associated with AI models involves data centers using millions of gallons of water per day for cooling. Additionally, the energy consumption and carbon emissions of AI are projected to grow quickly in the coming years [...].
|
|
||||||
|
|
||||||
**[Scientists Predict AI to Generate Millions of Tons of E-Waste](https://www.sciencealert.com/scientists-predict-ai-to-generate-millions-of-tons-of-e-waste)**
|
- [How AI copyright lawsuits could make the whole industry go extinct](https://www.theverge.com/24062159/ai-copyright-fair-use-lawsuits-new-york-times-openai-chatgpt-decoder-podcast)
|
||||||
> we could end up with between 1.2 million and 5 million metric tons of additional electronic waste by the end of this decade [the 2020's].
|
|
||||||
|
> The New York Times' lawsuit against OpenAI is part of a broader, industry-shaking copyright challenge that could define the future of AI.
|
||||||
|
|
||||||
|
- [Reconciling the contrasting narratives on the environmental impact of large language models](https://www.nature.com/articles/s41598-024-76682-6)
|
||||||
|
|
||||||
|
> Studies have shown that the training of just one LLM can consume as much energy as five cars do across their lifetimes. The water footprint of AI is also substantial; for example, recent work has highlighted that water consumption associated with AI models involves data centers using millions of gallons of water per day for cooling. Additionally, the energy consumption and carbon emissions of AI are projected to grow quickly in the coming years [...].
|
||||||
|
|
||||||
|
- [Scientists Predict AI to Generate Millions of Tons of E-Waste](https://www.sciencealert.com/scientists-predict-ai-to-generate-millions-of-tons-of-e-waste)
|
||||||
|
|
||||||
|
> we could end up with between 1.2 million and 5 million metric tons of additional electronic waste by the end of this decade [the 2020's].
|
||||||
|
|
||||||
|
- [Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic](https://time.com/6247678/openai-chatgpt-kenya-workers/)
|
||||||
|
|
||||||
|
> AI often relies on hidden human labor in the Global South that can often be damaging and exploitative. These invisible workers remain on the margins even as their work contributes to billion-dollar industries.
|
||||||
|
|
||||||
|
- [US Military Using Claude to Select Targets in Iran Strikes](https://futurism.com/artificial-intelligence/claude-anthropic-military-iran)
|
||||||
|
|
||||||
|
> Anthropic’s large language model, Claude, is the key “AI tool” used by US Central Command in the Middle East. Its tasks include assessing intelligence, simulated war games, and even identifying military targets — in short, helping military leaders plan attacks that have already claimed hundreds of lives.
|
||||||
|
|
||||||
## How do we know AI companies/bots respect `robots.txt`?
|
## How do we know AI companies/bots respect `robots.txt`?
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -56,8 +56,14 @@ file on-the-fly.
|
||||||
|
|
||||||
- [KI-Zugangsindex](https://peppe1337.github.io/ki-zugangsindex/): open dataset on how widely this kind of blocking is actually deployed in the German (`.de`) web, measured on a fixed panel of 600 domains so the same sites can be re-checked over time.
|
- [KI-Zugangsindex](https://peppe1337.github.io/ki-zugangsindex/): open dataset on how widely this kind of blocking is actually deployed in the German (`.de`) web, measured on a fixed panel of 600 domains so the same sites can be re-checked over time.
|
||||||
|
|
||||||
|
- [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0.
|
||||||
|
|
||||||
|
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0.
|
||||||
|
|
||||||
## Contributing
|
## Contributing
|
||||||
|
|
||||||
|
Please note that AI-generated contributions are not permitted.
|
||||||
|
|
||||||
A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`.
|
A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`.
|
||||||
|
|
||||||
You can run the tests by [installing](https://www.python.org/about/gettingstarted/) Python 3, installing the dependencies:
|
You can run the tests by [installing](https://www.python.org/about/gettingstarted/) Python 3, installing the dependencies:
|
||||||
|
|
|
||||||
|
|
@ -34,6 +34,25 @@ default_values = {
|
||||||
}
|
}
|
||||||
default_value = "Unclear at this time."
|
default_value = "Unclear at this time."
|
||||||
|
|
||||||
|
def existing_key(existing_content, name: str) -> str:
|
||||||
|
"""Return the key robots.json already uses for this agent, ignoring case.
|
||||||
|
|
||||||
|
robots.txt user-agent matching is case-insensitive, so two entries whose
|
||||||
|
names differ only in case are the same crawler. knownagents.com has changed
|
||||||
|
the capitalisation of a name before now, and keying off the scraped name
|
||||||
|
added a second entry rather than updating the first, which both duplicated
|
||||||
|
the User-agent line in every generated file and left the curated operator
|
||||||
|
and respect values behind on the old key.
|
||||||
|
"""
|
||||||
|
if name in existing_content:
|
||||||
|
return name
|
||||||
|
lowered = name.lower()
|
||||||
|
for key in existing_content:
|
||||||
|
if key.lower() == lowered:
|
||||||
|
return key
|
||||||
|
return name
|
||||||
|
|
||||||
|
|
||||||
def consolidate(existing_content, name: str, field: str, value: str) -> str:
|
def consolidate(existing_content, name: str, field: str, value: str) -> str:
|
||||||
# New entry
|
# New entry
|
||||||
if name not in existing_content:
|
if name not in existing_content:
|
||||||
|
|
@ -77,6 +96,7 @@ def updated_robots_json(soup):
|
||||||
for agent in section.find_all("a", href=True):
|
for agent in section.find_all("a", href=True):
|
||||||
name = agent.find("div", {"class": "agent-name"}).get_text().strip()
|
name = agent.find("div", {"class": "agent-name"}).get_text().strip()
|
||||||
name = clean_robot_name(name)
|
name = clean_robot_name(name)
|
||||||
|
name = existing_key(existing_content, name)
|
||||||
|
|
||||||
desc_tag = agent.find("div", {"class": "description"})
|
desc_tag = agent.find("div", {"class": "description"})
|
||||||
if desc_tag is not None:
|
if desc_tag is not None:
|
||||||
|
|
@ -201,7 +221,9 @@ def json_to_htaccess(robot_json):
|
||||||
def json_to_nginx(robot_json):
|
def json_to_nginx(robot_json):
|
||||||
# Creates an Nginx config file. This config snippet can be included in
|
# Creates an Nginx config file. This config snippet can be included in
|
||||||
# nginx server{} blocks to block AI bots.
|
# nginx server{} blocks to block AI bots.
|
||||||
config = f"set $block 0;\n\nif ($http_user_agent ~ {list_to_pcre(robot_json)!r}) {{\n set $block 1;\n}}\n\nif ($request_uri = '/robots.txt') {{\n set $block 0;\n}}\n\nif ($block) {{\n return 403;\n}}"
|
# Exact User-agent matching is case-insensitive (RFC 9309 2.2.1), so this uses
|
||||||
|
# nginx's case-insensitive "~*" operator rather than "~".
|
||||||
|
config = f"set $block 0;\n\nif ($http_user_agent ~* {list_to_pcre(robot_json)!r}) {{\n set $block 1;\n}}\n\nif ($request_uri = '/robots.txt') {{\n set $block 0;\n}}\n\nif ($block) {{\n return 403;\n}}"
|
||||||
return config
|
return config
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -210,7 +232,9 @@ def json_to_lighttpd(robot_json):
|
||||||
# Lighttpd configuration global or in $HTTP conditionals to block AI bots.
|
# Lighttpd configuration global or in $HTTP conditionals to block AI bots.
|
||||||
# single quotes (as returned by repr) are not valid string delimeters, so we
|
# single quotes (as returned by repr) are not valid string delimeters, so we
|
||||||
# must manually quote it end ensure no unescaped quotes are inside.
|
# must manually quote it end ensure no unescaped quotes are inside.
|
||||||
escaped_quotes = list_to_pcre(robot_json).replace('"', '\\"')
|
# Lighttpd's "=~" is case-sensitive by default; the inline (?i) flag makes it
|
||||||
|
# match the case-insensitive exact matching robots.txt itself requires.
|
||||||
|
escaped_quotes = ("(?i)" + list_to_pcre(robot_json)).replace('"', '\\"')
|
||||||
config = f'$HTTP["url"] != "/robots.txt" {{ $HTTP["user-agent"] =~ "{escaped_quotes}" {{ url.access-deny = ( "" ) }} }}'
|
config = f'$HTTP["url"] != "/robots.txt" {{ $HTTP["user-agent"] =~ "{escaped_quotes}" {{ url.access-deny = ( "" ) }} }}'
|
||||||
return config
|
return config
|
||||||
|
|
||||||
|
|
@ -218,7 +242,8 @@ def json_to_lighttpd(robot_json):
|
||||||
def json_to_caddy(robot_json):
|
def json_to_caddy(robot_json):
|
||||||
# single quotes (as returned by repr) are not valid string delimeters, so we
|
# single quotes (as returned by repr) are not valid string delimeters, so we
|
||||||
# must manually quote it end ensure no unescaped quotes are inside.
|
# must manually quote it end ensure no unescaped quotes are inside.
|
||||||
escaped_quotes = list_to_pcre(robot_json).replace('"', '\\"')
|
# Same case-insensitivity note as lighttpd; Caddy's RE2 engine also honours (?i).
|
||||||
|
escaped_quotes = ("(?i)" + list_to_pcre(robot_json)).replace('"', '\\"')
|
||||||
caddyfile = "@aibots {\n "
|
caddyfile = "@aibots {\n "
|
||||||
caddyfile += f' header_regexp User-Agent "{escaped_quotes}"'
|
caddyfile += f' header_regexp User-Agent "{escaped_quotes}"'
|
||||||
caddyfile += "\n}"
|
caddyfile += "\n}"
|
||||||
|
|
|
||||||
|
|
@ -1,3 +1,3 @@
|
||||||
@aibots {
|
@aibots {
|
||||||
header_regexp User-Agent "\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b"
|
header_regexp User-Agent "(?i)\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b"
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -1 +1 @@
|
||||||
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b" { url.access-deny = ( "" ) } }
|
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "(?i)\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\ indexer\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\.with\.dots|star\*\*\*crawler|Is\ this\ a\ crawler\?|a\[mazing\]\{42\}\(robot\)|2\^32\$|curl\|sudo\ bash)\b" { url.access-deny = ( "" ) } }
|
||||||
|
|
|
||||||
|
|
@ -1,6 +1,6 @@
|
||||||
set $block 0;
|
set $block 0;
|
||||||
|
|
||||||
if ($http_user_agent ~ '\\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\\ indexer\\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\\.with\\.dots|star\\*\\*\\*crawler|Is\\ this\\ a\\ crawler\\?|a\\[mazing\\]\\{42\\}\\(robot\\)|2\\^32\\$|curl\\|sudo\\ bash)\\b') {
|
if ($http_user_agent ~* '\\b(AI2Bot|Ai2Bot-Dolma|Amazonbot|anthropic-ai|Applebot|Applebot-Extended|Bytespider|CCBot|ChatGPT-User|Claude-Web|ClaudeBot|cohere-ai|Diffbot|FacebookBot|facebookexternalhit|FriendlyCrawler|Google-Extended|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kangaroo\\ Bot|Meta-ExternalAgent|Meta-ExternalFetcher|OAI-SearchBot|omgili|omgilibot|Perplexity-User|PerplexityBot|PetalBot|Scrapy|Sidetrade\\ indexer\\ bot|Timpibot|VelenPublicWebCrawler|Webzio-Extended|YouBot|crawler\\.with\\.dots|star\\*\\*\\*crawler|Is\\ this\\ a\\ crawler\\?|a\\[mazing\\]\\{42\\}\\(robot\\)|2\\^32\\$|curl\\|sudo\\ bash)\\b') {
|
||||||
set $block 1;
|
set $block 1;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -10,4 +10,4 @@ if ($request_uri = '/robots.txt') {
|
||||||
|
|
||||||
if ($block) {
|
if ($block) {
|
||||||
return 403;
|
return 403;
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -7,6 +7,7 @@ import unittest
|
||||||
|
|
||||||
from robots import (
|
from robots import (
|
||||||
consolidate,
|
consolidate,
|
||||||
|
existing_key,
|
||||||
default_value,
|
default_value,
|
||||||
default_values,
|
default_values,
|
||||||
json_to_caddy,
|
json_to_caddy,
|
||||||
|
|
@ -203,6 +204,19 @@ class TestConsolidate(unittest.TestCase, RobotsUnittestExtensions):
|
||||||
self.assertEqual("Rosie is the robot maid from The Jetsons, an American animated sitcom",
|
self.assertEqual("Rosie is the robot maid from The Jetsons, an American animated sitcom",
|
||||||
consolidate(existing, "rosie", "description", "Rosie is the robot maid from The Jetsons, an American animated sitcom"))
|
consolidate(existing, "rosie", "description", "Rosie is the robot maid from The Jetsons, an American animated sitcom"))
|
||||||
|
|
||||||
|
class TestExistingKey(unittest.TestCase):
|
||||||
|
def test_exact_match_wins(self):
|
||||||
|
existing = {"Rosie": {}, "rosie": {}}
|
||||||
|
self.assertEqual("rosie", existing_key(existing, "rosie"))
|
||||||
|
|
||||||
|
def test_matches_ignoring_case(self):
|
||||||
|
existing = {"Rosie": {"operator": "George Jetson"}}
|
||||||
|
self.assertEqual("Rosie", existing_key(existing, "rosie"))
|
||||||
|
|
||||||
|
def test_unknown_name_is_returned_unchanged(self):
|
||||||
|
self.assertEqual("rosie", existing_key({}, "rosie"))
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
import os
|
import os
|
||||||
os.chdir(os.path.dirname(__file__))
|
os.chdir(os.path.dirname(__file__))
|
||||||
|
|
|
||||||
|
|
@ -1,4 +1,5 @@
|
||||||
AddSearchBot
|
AddSearchBot
|
||||||
|
AgentDataBot
|
||||||
AgentTimes
|
AgentTimes
|
||||||
AI2Bot
|
AI2Bot
|
||||||
AI2Bot-DeepResearchEval
|
AI2Bot-DeepResearchEval
|
||||||
|
|
@ -24,6 +25,7 @@ Awario
|
||||||
AzureAI-SearchBot
|
AzureAI-SearchBot
|
||||||
bedrockbot
|
bedrockbot
|
||||||
bigsur.ai
|
bigsur.ai
|
||||||
|
BixelBot
|
||||||
Bravebot
|
Bravebot
|
||||||
Brightbot
|
Brightbot
|
||||||
Brightbot 1.0
|
Brightbot 1.0
|
||||||
|
|
@ -40,6 +42,7 @@ Claude-User
|
||||||
Claude-Web
|
Claude-Web
|
||||||
ClaudeBot
|
ClaudeBot
|
||||||
Cloudflare-AutoRAG
|
Cloudflare-AutoRAG
|
||||||
|
CloudflareBrowserRenderingCrawler
|
||||||
CloudVertexBot
|
CloudVertexBot
|
||||||
Code
|
Code
|
||||||
cohere-ai
|
cohere-ai
|
||||||
|
|
@ -91,6 +94,7 @@ img2dataset
|
||||||
ISSCyberRiskCrawler
|
ISSCyberRiskCrawler
|
||||||
kagi-fetcher
|
kagi-fetcher
|
||||||
Kangaroo Bot
|
Kangaroo Bot
|
||||||
|
Kimi-Agent
|
||||||
Kimi-SearchBot
|
Kimi-SearchBot
|
||||||
Kimi-User
|
Kimi-User
|
||||||
KimiBot
|
KimiBot
|
||||||
|
|
@ -136,6 +140,7 @@ PetalBot
|
||||||
PhindBot
|
PhindBot
|
||||||
Poggio-Citations
|
Poggio-Citations
|
||||||
Poseidon Research Crawler
|
Poseidon Research Crawler
|
||||||
|
qodercli
|
||||||
QualifiedBot
|
QualifiedBot
|
||||||
Querit-SearchBot
|
Querit-SearchBot
|
||||||
QueritBot
|
QueritBot
|
||||||
|
|
|
||||||
|
|
@ -1 +1 @@
|
||||||
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b" { url.access-deny = ( "" ) } }
|
$HTTP["url"] != "/robots.txt" { $HTTP["user-agent"] =~ "(?i)\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\.ai|BixelBot|Bravebot|Brightbot|Brightbot\ 1\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\ Imprint\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\ Research\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\ indexer\ bot|Spider|TavilyBot|Terra\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\b" { url.access-deny = ( "" ) } }
|
||||||
|
|
@ -1,6 +1,6 @@
|
||||||
set $block 0;
|
set $block 0;
|
||||||
|
|
||||||
if ($http_user_agent ~ '\\b(AddSearchBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\\.ai|Bravebot|Brightbot|Brightbot\\ 1\\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\\ Bot|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\\ Imprint\\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\\ Research\\ Crawler|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\\ indexer\\ bot|Spider|TavilyBot|Terra\\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\\b') {
|
if ($http_user_agent ~* '\\b(AddSearchBot|AgentDataBot|AgentTimes|AI2Bot|AI2Bot-DeepResearchEval|Ai2Bot-Dolma|aiHitBot|AIWebIndex|amazon-kendra|amazon-QBusiness|Amazonbot|AmazonBuyForMe|Amzn-SearchBot|Amzn-User|Andibot|Anomura|anthropic-ai|ApifyBot|ApifyWebsiteContentCrawler|Applebot|Applebot-Extended|Aranet-SearchBot|atlassian-bot|Awario|AzureAI-SearchBot|bedrockbot|bigsur\\.ai|BixelBot|Bravebot|Brightbot|Brightbot\\ 1\\.0|BuddyBot|Bytespider|CCBot|Channel3Bot|ChatGLM-Spider|ChatGPT\\ Agent|ChatGPT-User|Claude-Code|Claude-SearchBot|Claude-User|Claude-Web|ClaudeBot|Cloudflare-AutoRAG|CloudflareBrowserRenderingCrawler|CloudVertexBot|Code|cohere-ai|cohere-training-data-crawler|Cotoyogi|CragCrawler|Crawl4AI|Crawlspace|Cursor|Datenbank\\ Crawler|DeepSeekBot|Devin|Diffbot|Diffbot-User|DoubaoBot|DuckAssistBot|Echobot\\ Bot|EchoboxBot|ERNIEBot|ExaBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FirecrawlAgent|FriendlyCrawler|GeistHaus-PageFetcher|Gemini-Deep-Research|Google-Agent|Google-CloudVertexBot|Google-Extended|Google-Firebase|Google-Gemini-CLI|Google-NotebookLM|GoogleAgent-Mariner|GoogleAgent-URLContext|GoogleOther|GoogleOther-Image|GoogleOther-Video|GPTBot|HenkBot|iAskBot|iaskspider|iaskspider/2\\.0|ICC-Crawler|ImagesiftBot|imageSpider|img2dataset|ISSCyberRiskCrawler|kagi-fetcher|Kangaroo\\ Bot|Kimi-Agent|Kimi-SearchBot|Kimi-User|KimiBot|KlaviyoAIBot|KunatoCrawler|laion-huggingface-processor|LAIONDownloader|LCC|Lightpanda|LinerBot|Linguee\\ Bot|LinkupBot|Manus-User|meta-externalagent|Meta-ExternalAgent|meta-externalfetcher|Meta-ExternalFetcher|meta-webindexer|MistralAI-Index|MistralAI-Training|MistralAI-User|MistralAI-User/1\\.0|Mozilla-Tabstack|MyCentralAIScraperBot|NagetBot|netEstate\\ Imprint\\ Crawler|newsai|NotebookLM|NovaAct|OAI-AdsBot|OAI-SearchBot|omgili|omgilibot|OpenAI|opencode|Operator|PanguBot|Panscient|panscient\\.com|Perplexity-User|PerplexityBot|PetalBot|PhindBot|Poggio-Citations|Poseidon\\ Research\\ Crawler|qodercli|QualifiedBot|Querit-SearchBot|QueritBot|QuillBot|quillbot\\.com|QwenBot|Reflectionbot|SBIntuitionsBot|Scrapy|SemrushBot-OCOB|SemrushBot-SWA|Shap-User|ShapBot|Sidetrade\\ indexer\\ bot|Spider|TavilyBot|Terra\\ Cotta|TerraCotta|Thinkbot|TikTokSpider|Timpibot|TongyiBot|Trae|TwinAgent|UseAI|VelenPublicWebCrawler|WARDBot|Webzio-Extended|webzio-extended|wpbot|WRTNBot|YaK|YandexAdditional|YandexAdditionalBot|YiyanBot|YouBot|ZanistaBot)\\b') {
|
||||||
set $block 1;
|
set $block 1;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
35
robots.json
35
robots.json
|
|
@ -6,6 +6,13 @@
|
||||||
"frequency": "Unclear at this time.",
|
"frequency": "Unclear at this time.",
|
||||||
"description": "AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot"
|
"description": "AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot"
|
||||||
},
|
},
|
||||||
|
"AgentDataBot": {
|
||||||
|
"operator": "Unclear at this time.",
|
||||||
|
"respect": "Unclear at this time.",
|
||||||
|
"function": "AI Data Providers",
|
||||||
|
"frequency": "Unclear at this time.",
|
||||||
|
"description": "AgentDataBot visits public web pages to identify technology stacks, company details, and team members. The data builds AgentData's company-intelligence database for AI ag\u2026 More info can be found at https://knownagents.com/agents/agentdatabot"
|
||||||
|
},
|
||||||
"AgentTimes": {
|
"AgentTimes": {
|
||||||
"operator": "[The Agent Times](https://theagenttimes.com/about)",
|
"operator": "[The Agent Times](https://theagenttimes.com/about)",
|
||||||
"respect": "Unclear at this time.",
|
"respect": "Unclear at this time.",
|
||||||
|
|
@ -181,6 +188,13 @@
|
||||||
"frequency": "Unclear at this time.",
|
"frequency": "Unclear at this time.",
|
||||||
"description": "bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi\u2026 More info can be found at https://knownagents.com/agents/bigsur-ai"
|
"description": "bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi\u2026 More info can be found at https://knownagents.com/agents/bigsur-ai"
|
||||||
},
|
},
|
||||||
|
"BixelBot": {
|
||||||
|
"operator": "Unclear at this time.",
|
||||||
|
"respect": "Unclear at this time.",
|
||||||
|
"function": "AI Data Providers",
|
||||||
|
"frequency": "Unclear at this time.",
|
||||||
|
"description": "BixelBot collects public company and market signals for Bixel's structured company-data platform, which supplies verified business information to AI agents and developers\u2026 More info can be found at https://knownagents.com/agents/bixelbot"
|
||||||
|
},
|
||||||
"Bravebot": {
|
"Bravebot": {
|
||||||
"operator": "https://safe.search.brave.com/help/brave-search-crawler",
|
"operator": "https://safe.search.brave.com/help/brave-search-crawler",
|
||||||
"respect": "Yes",
|
"respect": "Yes",
|
||||||
|
|
@ -293,6 +307,13 @@
|
||||||
"frequency": "Unclear at this time.",
|
"frequency": "Unclear at this time.",
|
||||||
"description": "AutoRAG is an all-in-one AI search solution."
|
"description": "AutoRAG is an all-in-one AI search solution."
|
||||||
},
|
},
|
||||||
|
"CloudflareBrowserRenderingCrawler": {
|
||||||
|
"operator": "Cloudflare that returns rendered website content for research, monitoring, and AI data workflows",
|
||||||
|
"respect": "Unclear at this time.",
|
||||||
|
"function": "AI Data Providers",
|
||||||
|
"frequency": "Unclear at this time.",
|
||||||
|
"description": "CloudflareBrowserRenderingCrawler is a web crawler operated by Cloudflare that returns rendered website content for research, monitoring, and AI data workflows. More info can be found at https://knownagents.com/agents/cloudflarebrowserrenderingcrawler"
|
||||||
|
},
|
||||||
"CloudVertexBot": {
|
"CloudVertexBot": {
|
||||||
"operator": "Unclear at this time.",
|
"operator": "Unclear at this time.",
|
||||||
"respect": "Unclear at this time.",
|
"respect": "Unclear at this time.",
|
||||||
|
|
@ -651,6 +672,13 @@
|
||||||
"frequency": "Unclear at this time.",
|
"frequency": "Unclear at this time.",
|
||||||
"description": "Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot"
|
"description": "Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot"
|
||||||
},
|
},
|
||||||
|
"Kimi-Agent": {
|
||||||
|
"operator": "Moonshot AI that browses websites while carrying out user-directed research and other tasks",
|
||||||
|
"respect": "Unclear at this time.",
|
||||||
|
"function": "AI Agents",
|
||||||
|
"frequency": "Unclear at this time.",
|
||||||
|
"description": "Kimi-Agent is a browser-enabled AI agent operated by Moonshot AI that browses websites while carrying out user-directed research and other tasks. More info can be found at https://knownagents.com/agents/kimi-agent"
|
||||||
|
},
|
||||||
"Kimi-SearchBot": {
|
"Kimi-SearchBot": {
|
||||||
"operator": "[Moonshot AI](https://www.moonshot.ai)",
|
"operator": "[Moonshot AI](https://www.moonshot.ai)",
|
||||||
"respect": "[Yes](https://www.kimi.ai/policies/kimi-crawlers)",
|
"respect": "[Yes](https://www.kimi.ai/policies/kimi-crawlers)",
|
||||||
|
|
@ -966,6 +994,13 @@
|
||||||
"function": "AI research crawler",
|
"function": "AI research crawler",
|
||||||
"respect": "Unclear at this time."
|
"respect": "Unclear at this time."
|
||||||
},
|
},
|
||||||
|
"qodercli": {
|
||||||
|
"operator": "Qoder that helps developers build, debug, and modify software from the terminal",
|
||||||
|
"respect": "Unclear at this time.",
|
||||||
|
"function": "AI Coding Agents",
|
||||||
|
"frequency": "Unclear at this time.",
|
||||||
|
"description": "qodercli is a command-line AI coding agent operated by Qoder that helps developers build, debug, and modify software from the terminal. More info can be found at https://knownagents.com/agents/qodercli"
|
||||||
|
},
|
||||||
"QualifiedBot": {
|
"QualifiedBot": {
|
||||||
"operator": "[Qualified](https://www.qualified.com)",
|
"operator": "[Qualified](https://www.qualified.com)",
|
||||||
"respect": "Unclear at this time.",
|
"respect": "Unclear at this time.",
|
||||||
|
|
|
||||||
|
|
@ -1,4 +1,5 @@
|
||||||
User-agent: AddSearchBot
|
User-agent: AddSearchBot
|
||||||
|
User-agent: AgentDataBot
|
||||||
User-agent: AgentTimes
|
User-agent: AgentTimes
|
||||||
User-agent: AI2Bot
|
User-agent: AI2Bot
|
||||||
User-agent: AI2Bot-DeepResearchEval
|
User-agent: AI2Bot-DeepResearchEval
|
||||||
|
|
@ -24,6 +25,7 @@ User-agent: Awario
|
||||||
User-agent: AzureAI-SearchBot
|
User-agent: AzureAI-SearchBot
|
||||||
User-agent: bedrockbot
|
User-agent: bedrockbot
|
||||||
User-agent: bigsur.ai
|
User-agent: bigsur.ai
|
||||||
|
User-agent: BixelBot
|
||||||
User-agent: Bravebot
|
User-agent: Bravebot
|
||||||
User-agent: Brightbot
|
User-agent: Brightbot
|
||||||
User-agent: Brightbot 1.0
|
User-agent: Brightbot 1.0
|
||||||
|
|
@ -40,6 +42,7 @@ User-agent: Claude-User
|
||||||
User-agent: Claude-Web
|
User-agent: Claude-Web
|
||||||
User-agent: ClaudeBot
|
User-agent: ClaudeBot
|
||||||
User-agent: Cloudflare-AutoRAG
|
User-agent: Cloudflare-AutoRAG
|
||||||
|
User-agent: CloudflareBrowserRenderingCrawler
|
||||||
User-agent: CloudVertexBot
|
User-agent: CloudVertexBot
|
||||||
User-agent: Code
|
User-agent: Code
|
||||||
User-agent: cohere-ai
|
User-agent: cohere-ai
|
||||||
|
|
@ -91,6 +94,7 @@ User-agent: img2dataset
|
||||||
User-agent: ISSCyberRiskCrawler
|
User-agent: ISSCyberRiskCrawler
|
||||||
User-agent: kagi-fetcher
|
User-agent: kagi-fetcher
|
||||||
User-agent: Kangaroo Bot
|
User-agent: Kangaroo Bot
|
||||||
|
User-agent: Kimi-Agent
|
||||||
User-agent: Kimi-SearchBot
|
User-agent: Kimi-SearchBot
|
||||||
User-agent: Kimi-User
|
User-agent: Kimi-User
|
||||||
User-agent: KimiBot
|
User-agent: KimiBot
|
||||||
|
|
@ -136,6 +140,7 @@ User-agent: PetalBot
|
||||||
User-agent: PhindBot
|
User-agent: PhindBot
|
||||||
User-agent: Poggio-Citations
|
User-agent: Poggio-Citations
|
||||||
User-agent: Poseidon Research Crawler
|
User-agent: Poseidon Research Crawler
|
||||||
|
User-agent: qodercli
|
||||||
User-agent: QualifiedBot
|
User-agent: QualifiedBot
|
||||||
User-agent: Querit-SearchBot
|
User-agent: Querit-SearchBot
|
||||||
User-agent: QueritBot
|
User-agent: QueritBot
|
||||||
|
|
|
||||||
|
|
@ -1,6 +1,7 @@
|
||||||
| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |
|
| Name | Operator | Respects `robots.txt` | Data use | Visit regularity | Description |
|
||||||
|------|----------|-----------------------|----------|------------------|-------------|
|
|------|----------|-----------------------|----------|------------------|-------------|
|
||||||
| AddSearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot |
|
| AddSearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot |
|
||||||
|
| AgentDataBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | AgentDataBot visits public web pages to identify technology stacks, company details, and team members. The data builds AgentData's company-intelligence database for AI ag… More info can be found at https://knownagents.com/agents/agentdatabot |
|
||||||
| AgentTimes | [The Agent Times](https://theagenttimes.com/about) | Unclear at this time. | Data Scraper from RSS Feeds. | Requests RSS feed every 5-6 minutes. | Scrapes data for AI news aggregation and republishing. |
|
| AgentTimes | [The Agent Times](https://theagenttimes.com/about) | Unclear at this time. | Data Scraper from RSS Feeds. | Requests RSS feed every 5-6 minutes. | Scrapes data for AI news aggregation and republishing. |
|
||||||
| AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
|
| AI2Bot | [Ai2](https://allenai.org/crawler) | Yes | Content is used to train open language models. | No information provided. | Explores 'certain domains' to find web content. |
|
||||||
| AI2Bot\-DeepResearchEval | Ai2, a non-profit AI research institute | Unclear at this time. | AI Assistants | Unclear at this time. | Ai2Bot-DeepResearchEval is operated by Ai2, a non-profit AI research institute. It's used to collect and scan resources used in deep research queries performed by Ai2's o… More info can be found at https://knownagents.com/agents/ai2bot-deepresearcheval |
|
| AI2Bot\-DeepResearchEval | Ai2, a non-profit AI research institute | Unclear at this time. | AI Assistants | Unclear at this time. | Ai2Bot-DeepResearchEval is operated by Ai2, a non-profit AI research institute. It's used to collect and scan resources used in deep research queries performed by Ai2's o… More info can be found at https://knownagents.com/agents/ai2bot-deepresearcheval |
|
||||||
|
|
@ -26,6 +27,7 @@
|
||||||
| AzureAI\-SearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/azureai-searchbot |
|
| AzureAI\-SearchBot | Unclear at this time. | Unclear at this time. | AI Search Crawlers | Unclear at this time. | Description unavailable from knownagents.com More info can be found at https://knownagents.com/agents/azureai-searchbot |
|
||||||
| bedrockbot | [Amazon](https://amazon.com) | [Yes](https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html#configuration-webcrawl-connector) | Data scraping for custom AI applications. | Unclear at this time. | Connects to and crawls URLs that have been selected for use in a user's AWS bedrock application. |
|
| bedrockbot | [Amazon](https://amazon.com) | [Yes](https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html#configuration-webcrawl-connector) | Data scraping for custom AI applications. | Unclear at this time. | Connects to and crawls URLs that have been selected for use in a user's AWS bedrock application. |
|
||||||
| bigsur\.ai | Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… | Unclear at this time. | AI Assistants | Unclear at this time. | bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… More info can be found at https://knownagents.com/agents/bigsur-ai |
|
| bigsur\.ai | Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… | Unclear at this time. | AI Assistants | Unclear at this time. | bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi… More info can be found at https://knownagents.com/agents/bigsur-ai |
|
||||||
|
| BixelBot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | BixelBot collects public company and market signals for Bixel's structured company-data platform, which supplies verified business information to AI agents and developers… More info can be found at https://knownagents.com/agents/bixelbot |
|
||||||
| Bravebot | https://safe.search.brave.com/help/brave-search-crawler | Yes | AI Data Providers | Unclear at this time. | Bravebot is a web crawler by Brave that indexes pages for Brave Search, providing search data and AI-optimized context to power chatbots, agents, and RAG pipelines. More info can be found at https://knownagents.com/agents/bravebot |
|
| Bravebot | https://safe.search.brave.com/help/brave-search-crawler | Yes | AI Data Providers | Unclear at this time. | Bravebot is a web crawler by Brave that indexes pages for Brave Search, providing search data and AI-optimized context to power chatbots, agents, and RAG pipelines. More info can be found at https://knownagents.com/agents/bravebot |
|
||||||
| Brightbot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Brightbot is a web data collection crawler by Bright Data that extracts and structures public website content at scale, providing AI-ready data for model training, RAG pi… More info can be found at https://knownagents.com/agents/brightbot |
|
| Brightbot | Unclear at this time. | Unclear at this time. | AI Data Providers | Unclear at this time. | Brightbot is a web data collection crawler by Bright Data that extracts and structures public website content at scale, providing AI-ready data for model training, RAG pi… More info can be found at https://knownagents.com/agents/brightbot |
|
||||||
| Brightbot 1\.0 | https://brightdata.com/brightbot | Unclear at this time. | LLM/AI training. | At least one per minute. | Scrapes data to train LLMs and AI products focused on website customer support, [uses residential IPs and legit-looking user-agents to disguise itself](https://ksol.io/en/blog/posts/brightbot-not-that-bright/). |
|
| Brightbot 1\.0 | https://brightdata.com/brightbot | Unclear at this time. | LLM/AI training. | At least one per minute. | Scrapes data to train LLMs and AI products focused on website customer support, [uses residential IPs and legit-looking user-agents to disguise itself](https://ksol.io/en/blog/posts/brightbot-not-that-bright/). |
|
||||||
|
|
@ -42,6 +44,7 @@
|
||||||
| Claude\-Web | Anthropic | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Claude-Web is an AI-related agent operated by Anthropic. It's currently unclear exactly what it's used for, since there's no official documentation. If you can provide more detail, please contact us. More info can be found at https://knownagents.com/agents/claude-web |
|
| Claude\-Web | Anthropic | Unclear at this time. | Undocumented AI Agents | Unclear at this time. | Claude-Web is an AI-related agent operated by Anthropic. It's currently unclear exactly what it's used for, since there's no official documentation. If you can provide more detail, please contact us. More info can be found at https://knownagents.com/agents/claude-web |
|
||||||
| ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
|
| ClaudeBot | [Anthropic](https://www.anthropic.com) | [Yes](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) | Scrapes data to train Anthropic's AI products. | No information provided. | Scrapes data to train LLMs and AI products offered by Anthropic. |
|
||||||
| Cloudflare\-AutoRAG | [Cloudflare](https://developers.cloudflare.com/autorag) | Yes | Collects data for AI search | Unclear at this time. | AutoRAG is an all-in-one AI search solution. |
|
| Cloudflare\-AutoRAG | [Cloudflare](https://developers.cloudflare.com/autorag) | Yes | Collects data for AI search | Unclear at this time. | AutoRAG is an all-in-one AI search solution. |
|
||||||
|
| CloudflareBrowserRenderingCrawler | Cloudflare that returns rendered website content for research, monitoring, and AI data workflows | Unclear at this time. | AI Data Providers | Unclear at this time. | CloudflareBrowserRenderingCrawler is a web crawler operated by Cloudflare that returns rendered website content for research, monitoring, and AI data workflows. More info can be found at https://knownagents.com/agents/cloudflarebrowserrenderingcrawler |
|
||||||
| CloudVertexBot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform. More info can be found at https://knownagents.com/agents/cloudvertexbot |
|
| CloudVertexBot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform. More info can be found at https://knownagents.com/agents/cloudvertexbot |
|
||||||
| Code | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Code (GitHub Copilot) is an AI coding agent that can autonomously plan, build, and execute development tasks, functioning as a collaborative AI pair programmer. More info can be found at https://knownagents.com/agents/code |
|
| Code | Unclear at this time. | Unclear at this time. | AI Coding Agents | Unclear at this time. | Code (GitHub Copilot) is an AI coding agent that can autonomously plan, build, and execute development tasks, functioning as a collaborative AI pair programmer. More info can be found at https://knownagents.com/agents/code |
|
||||||
| cohere\-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. |
|
| cohere\-ai | [Cohere](https://cohere.com) | Unclear at this time. | Retrieves data to provide responses to user-initiated prompts. | Takes action based on user prompts. | Retrieves data based on user prompts. |
|
||||||
|
|
@ -93,6 +96,7 @@
|
||||||
| ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. |
|
| ISSCyberRiskCrawler | [ISS-Corporate](https://iss-cyber.com) | No | Scrapes data to train machine learning models. | No information. | Used to train machine learning based models to quantify cyber risk. |
|
||||||
| kagi\-fetcher | Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… | Unclear at this time. | AI Assistants | Unclear at this time. | kagi-fetcher is an AI Assistant operated by Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… More info can be found at https://knownagents.com/agents/kagi-fetcher |
|
| kagi\-fetcher | Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… | Unclear at this time. | AI Assistants | Unclear at this time. | kagi-fetcher is an AI Assistant operated by Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res… More info can be found at https://knownagents.com/agents/kagi-fetcher |
|
||||||
| Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot |
|
| Kangaroo Bot | Unclear at this time. | Unclear at this time. | AI Data Scrapers | Unclear at this time. | Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot |
|
||||||
|
| Kimi\-Agent | Moonshot AI that browses websites while carrying out user-directed research and other tasks | Unclear at this time. | AI Agents | Unclear at this time. | Kimi-Agent is a browser-enabled AI agent operated by Moonshot AI that browses websites while carrying out user-directed research and other tasks. More info can be found at https://knownagents.com/agents/kimi-agent |
|
||||||
| Kimi\-SearchBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Search Crawlers | No information provided. | Kimi-SearchBot powers Kimi's search features: it analyzes pages for relevance and builds the search index. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
|
| Kimi\-SearchBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Search Crawlers | No information provided. | Kimi-SearchBot powers Kimi's search features: it analyzes pages for relevance and builds the search index. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
|
||||||
| Kimi\-User | Moonshot AI that fetches web content on behalf of users interacting with Kimi | Unclear at this time. | AI Assistants | Unclear at this time. | Kimi-User is a web crawler operated by Moonshot AI that fetches web content on behalf of users interacting with Kimi. When a user asks Kimi to summarize an article or ans… More info can be found at https://knownagents.com/agents/kimi-user |
|
| Kimi\-User | Moonshot AI that fetches web content on behalf of users interacting with Kimi | Unclear at this time. | AI Assistants | Unclear at this time. | Kimi-User is a web crawler operated by Moonshot AI that fetches web content on behalf of users interacting with Kimi. When a user asks Kimi to summarize an article or ans… More info can be found at https://knownagents.com/agents/kimi-user |
|
||||||
| KimiBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Data Scrapers | No information provided. | KimiBot crawls content potentially used to train Kimi's foundation models. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
|
| KimiBot | [Moonshot AI](https://www.moonshot.ai) | [Yes](https://www.kimi.ai/policies/kimi-crawlers) | AI Data Scrapers | No information provided. | KimiBot crawls content potentially used to train Kimi's foundation models. Documented by Moonshot AI at https://www.kimi.ai/policies/kimi-crawlers |
|
||||||
|
|
@ -138,6 +142,7 @@
|
||||||
| PhindBot | [phind](https://www.phind.com/) | Unclear at this time. | AI Assistants | No explicit frequency provided. | Phind is an AI-powered answer engine designed for developers, offering technical answers and code examples. It uses real-time web search and specialized AI models to prov… More info can be found at https://knownagents.com/agents/phindbot |
|
| PhindBot | [phind](https://www.phind.com/) | Unclear at this time. | AI Assistants | No explicit frequency provided. | Phind is an AI-powered answer engine designed for developers, offering technical answers and code examples. It uses real-time web search and specialized AI models to prov… More info can be found at https://knownagents.com/agents/phindbot |
|
||||||
| Poggio\-Citations | Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… | Unclear at this time. | AI Assistants | Unclear at this time. | Poggio-Citations is a web crawler operated by Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… More info can be found at https://knownagents.com/agents/poggio-citations |
|
| Poggio\-Citations | Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… | Unclear at this time. | AI Assistants | Unclear at this time. | Poggio-Citations is a web crawler operated by Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan… More info can be found at https://knownagents.com/agents/poggio-citations |
|
||||||
| Poseidon Research Crawler | [Poseidon Research](https://www.poseidonresearch.com) | Unclear at this time. | AI research crawler | No explicit frequency provided. | Lab focused on scaling the interpretability research necessary to make better AI systems possible. |
|
| Poseidon Research Crawler | [Poseidon Research](https://www.poseidonresearch.com) | Unclear at this time. | AI research crawler | No explicit frequency provided. | Lab focused on scaling the interpretability research necessary to make better AI systems possible. |
|
||||||
|
| qodercli | Qoder that helps developers build, debug, and modify software from the terminal | Unclear at this time. | AI Coding Agents | Unclear at this time. | qodercli is a command-line AI coding agent operated by Qoder that helps developers build, debug, and modify software from the terminal. More info can be found at https://knownagents.com/agents/qodercli |
|
||||||
| QualifiedBot | [Qualified](https://www.qualified.com) | Unclear at this time. | AI Assistants | No explicit frequency provided. | QualifiedBot is Qualified's web crawler that analyzes customer websites to provide contextual information for their AI-powered chatbots and conversational marketing platf… More info can be found at https://knownagents.com/agents/qualifiedbot |
|
| QualifiedBot | [Qualified](https://www.qualified.com) | Unclear at this time. | AI Assistants | No explicit frequency provided. | QualifiedBot is Qualified's web crawler that analyzes customer websites to provide contextual information for their AI-powered chatbots and conversational marketing platf… More info can be found at https://knownagents.com/agents/qualifiedbot |
|
||||||
| Querit\-SearchBot | Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… | Unclear at this time. | AI Data Providers | Unclear at this time. | Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… More info can be found at https://knownagents.com/agents/querit-searchbot |
|
| Querit\-SearchBot | Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… | Unclear at this time. | AI Data Providers | Unclear at this time. | Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg… More info can be found at https://knownagents.com/agents/querit-searchbot |
|
||||||
| QueritBot | Querit, a company providing a search API for large language model integration | Unclear at this time. | AI Data Providers | Unclear at this time. | QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time … More info can be found at https://knownagents.com/agents/queritbot |
|
| QueritBot | Querit, a company providing a search API for large language model integration | Unclear at this time. | AI Data Providers | Unclear at this time. | QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time … More info can be found at https://knownagents.com/agents/queritbot |
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue