"description":"AddSearchBot is a web crawler that indexes website content for AddSearch's AI-powered site search solution, collecting data to provide fast and accurate search results. More info can be found at https://knownagents.com/agents/addsearchbot"
"description":"Ai2Bot-DeepResearchEval is operated by Ai2, a non-profit AI research institute. It's used to collect and scan resources used in deep research queries performed by Ai2's o\u2026 More info can be found at https://knownagents.com/agents/ai2bot-deepresearcheval"
"operator":"Lyrenth that builds an AI-readable index of web content for AI systems",
"respect":"Unclear at this time.",
"function":"AI Data Providers",
"frequency":"Unclear at this time.",
"description":"AIWebIndex is a web crawler operated by Lyrenth that builds an AI-readable index of web content for AI systems. More info can be found at https://knownagents.com/agents/aiwebindex"
"description":"Amazon Kendra is a highly accurate intelligent search service that enables your users to search unstructured data using natural language. It returns specific answers to questions, giving users an experience that's close to interacting with a human expert. It is highly scalable and capable of meeting performance demands, tightly integrated with other AWS services such as Amazon S3 and Amazon Lex, and offers enterprise-grade security."
"description":"amazon-QBusiness is an Amazon Q Business web crawler that fetches and indexes web content for Amazon Q Business applications. More info can be found at https://knownagents.com/agents/amazon-qbusiness"
"description":"AmazonBuyForMe is an Amazon bot that crawls websites as part of the Amazon Buy For Me service. This bot visits product pages and e-commerce websites to gather product inf\u2026 More info can be found at https://knownagents.com/agents/amazonbuyforme"
"description":"Amzn-User is an AI assistant operated by Amazon, used for fetching web content to answer user queries through Alexa and other Amazon AI services. More info can be found at https://knownagents.com/agents/amzn-user"
"description":"ApifyBot is a web scraping and data extraction crawler by Apify that collects website content for use in AI, LLMs, RAG, and automation workflows. More info can be found at https://knownagents.com/agents/apifybot"
"description":"ApifyWebsiteContentCrawler is a web crawler by Apify that extracts and downloads full website content for use in AI, data analysis, and automation workflows. More info can be found at https://knownagents.com/agents/apifywebsitecontentcrawler"
"description":"Applebot is a web crawler used by Apple to index search results that allow the Siri AI Assistant to answer user questions. Siri's answers normally contain references to the website. More info can be found at https://knownagents.com/agents/applebot"
"function":"Powers features in Siri, Spotlight, Safari, Apple Intelligence, and others.",
"frequency":"Unclear at this time.",
"description":"Apple has a secondary user agent, Applebot-Extended ... [that is] used to train Apple's foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools."
"description":"atlassian-bot is a web crawler used to index website content for its AI search, assistants and agents available in its Rovo GenAI product."
"description":"Awario is an AI data scraper operated by Awario. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/awario"
"operator":"Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi\u2026",
"description":"bigsur.ai is a web crawler operated by Big Sur AI that fetches website content to enable AI-powered web agents, sales assistants, and content marketing solutions for busi\u2026 More info can be found at https://knownagents.com/agents/bigsur-ai"
"description":"Bravebot is a web crawler by Brave that indexes pages for Brave Search, providing search data and AI-optimized context to power chatbots, agents, and RAG pipelines. More info can be found at https://knownagents.com/agents/bravebot"
"description":"Brightbot is a web data collection crawler by Bright Data that extracts and structures public website content at scale, providing AI-ready data for model training, RAG pi\u2026 More info can be found at https://knownagents.com/agents/brightbot"
"description":"Scrapes data to train LLMs and AI products focused on website customer support, [uses residential IPs and legit-looking user-agents to disguise itself](https://ksol.io/en/blog/posts/brightbot-not-that-bright/)."
"description":"ChatGPT Agent is an AI agent created by OpenAI that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on behalf\u2026 More info can be found at https://knownagents.com/agents/chatgpt-agent"
"description":"ChatGPT-User is OpenAI's web crawler that visits websites when ChatGPT users request information. This enables ChatGPT to include links in its responses. More info can be found at https://knownagents.com/agents/chatgpt-user"
"description":"Claude Code is an AI coding agent by Anthropic that can build, debug, and ship code directly from the terminal, handling tasks like codebase onboarding, multi-file edits,\u2026 More info can be found at https://knownagents.com/agents/claude-code"
"function":"Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses.",
"description":"Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses."
"description":"Claude-User is dispatched by Anthropic's Claude AI assistant in response to user prompts, when it needs to fetch content to include in its answers. More info can be found at https://knownagents.com/agents/claude-user"
"description":"Claude-Web is an AI-related agent operated by Anthropic. It's currently unclear exactly what it's used for, since there's no official documentation. If you can provide more detail, please contact us. More info can be found at https://knownagents.com/agents/claude-web"
"description":"CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform. More info can be found at https://knownagents.com/agents/cloudvertexbot"
"description":"Code (GitHub Copilot) is an AI coding agent that can autonomously plan, build, and execute development tasks, functioning as a collaborative AI pair programmer. More info can be found at https://knownagents.com/agents/code",
"description":"cohere-training-data-crawler is a web crawler operated by Cohere to download training data for its LLMs (Large Language Models) that power its enterprise AI products. More info can be found at https://knownagents.com/agents/cohere-training-data-crawler"
"description":"CragCrawler is a web scraping bot operated by CragSoftware, a Brazil-based software company specializing in data engineering and AI web scraping services. The bot is used\u2026 More info can be found at https://knownagents.com/agents/cragcrawler"
"description":"Crawlspace is a web crawler platform that fetches and extracts website content for AI agents, RAG applications, and structured data workflows. More info can be found at https://knownagents.com/agents/crawlspace"
"description":"Cursor is an AI coding agent that helps write, edit, and understand code. More info can be found at https://knownagents.com/agents/cursor"
"description":"Datenbank Crawler is an AI data scraper operated by Datenbank. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/datenbank-crawler"
"description":"Devin is a software engineering AI assistant that can browse websites and perform web-based tasks, functioning as a collaborative AI teammate for engineering teams. More info can be found at https://knownagents.com/agents/devin"
"description":"Diffbot is a web crawler that extracts and structures website content using AI-powered visual understanding, providing knowledge graph data for applications like market i\u2026 More info can be found at https://knownagents.com/agents/diffbot"
"description":"DuckAssistBot is a web crawler that scans websites to collect content for DuckDuckGo's AI-assisted answers feature, which generates brief responses to search queries usin\u2026 More info can be found at https://knownagents.com/agents/duckassistbot"
"description":"Echobot Bot is an AI data scraper operated by Echobox. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/echobot-bot"
"description":"ExaBot is a web crawler that indexes web content to power Exa's AI search engine and semantic search APIs for AI applications. More info can be found at https://knownagents.com/agents/exabot"
"description":"Note that excluding FacebookExternalHit will block incorporating OpenGraph data when sharing in social media, including rich links in Apple's Messages app. [According to Meta](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/), its purpose is \"to crawl the content of an app or website that was shared on one of Meta\u2019s family of apps\u2026\". However, see discussions [here](https://github.com/ai-robots-txt/ai.robots.txt/pull/21) and [here](https://github.com/ai-robots-txt/ai.robots.txt/issues/40#issuecomment-2524591313) for evidence to the contrary."
"description":"FirecrawlAgent is a web crawler operated by Firecrawl that extracts web content and converts it into structured data for use in LLM and AI applications. More info can be found at https://knownagents.com/agents/firecrawlagent"
"description":"GeistHaus-PageFetcher is a web crawler operated by GeistHaus, a company developing AI systems for therapy and psychological assessment. This bot fetches web pages as part\u2026 More info can be found at https://knownagents.com/agents/geisthaus-pagefetcher"
"description":"Gemini-Deep-Research is the agent responsible for collecting and scanning resources used in Google Gemini's Deep Research feature, which acts as a personal research assis\u2026 More info can be found at https://knownagents.com/agents/gemini-deep-research"
"description":"Google-Agent is used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request. More info can be found at https://knownagents.com/agents/google-agent"
"description":"Gemini CLI is an AI coding agent by Google that can query and edit large codebases, generate apps from images or PDFs, and automate complex workflows directly from the te\u2026 More info can be found at https://knownagents.com/agents/google-gemini-cli"
"description":"Google-NotebookLM is an AI-powered research and note-taking assistant that helps users synthesize information from uploaded sources like documents, transcripts, or web co\u2026 More info can be found at https://knownagents.com/agents/google-notebooklm"
"description":"GoogleAgent-Mariner is an AI agent created by Google that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on \u2026 More info can be found at https://knownagents.com/agents/googleagent-mariner"
"description":"GoogleAgent-URLContext is a web fetcher operated by Google that retrieves web content on behalf of Gemini API users. When a developer provides a URL as context in a Gemin\u2026 More info can be found at https://knownagents.com/agents/googleagent-urlcontext"
"description":"\"Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development.\""
"description":"\"Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development.\"",
"description":"\"Used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development.\"",
"description":"Henkbot crawls the web on behalf of Valyu, an AI search infrastructure provider that indexes content for use in AI-powered retrieval pipelines. More info can be found at https://knownagents.com/agents/henkbot"
"function":"Scrapes data to train and support AI technologies.",
"frequency":"No information.",
"description":"Use the collected data for artificial intelligence technologies; provide data to third parties, including commercial companies; those companies can use the data for their own business."
"description":"Once images and text are downloaded from a webpage, ImageSift analyzes this data from the page and stores the information in an index. Their web intelligence products use this index to enable search and retrieval of similar images.",
"function":"ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support their suite of web intelligence products",
"description":"kagi-fetcher is an AI Assistant operated by Kagi that fetches web content to answer user queries through Kagi AI, their suite of AI-powered tools including Assistant, Res\u2026 More info can be found at https://knownagents.com/agents/kagi-fetcher"
"description":"Kangaroo Bot is used by the company Kangaroo LLM to download data to train AI models tailored to Australian language and culture. More info can be found at https://knownagents.com/agents/kangaroo-bot"
"description":"Kimi-User is a web crawler operated by Moonshot AI that fetches web content on behalf of users interacting with Kimi. When a user asks Kimi to summarize an article or ans\u2026 More info can be found at https://knownagents.com/agents/kimi-user"
"description":"KlaviyoAIBot is Klaviyo's web crawler that fetches publicly available pages from domains explicitly connected to user accounts to power the Kai Customer Agent feature. Th\u2026 More info can be found at https://knownagents.com/agents/klaviyoaibot"
"operator":"[Large-scale Artificial Intelligence Open Network](https://laion.ai/)",
"respect":"[No](https://laion.ai/faq/)",
"function":"AI tools and models for machine learning research.",
"frequency":"Unclear at this time.",
"description":"LAIONDownloader is a bot by LAION, a non-profit organization that provides datasets, tools and models to liberate machine learning research."
"description":"LinerBot is the web crawler used by Liner AI assistant to gather information from academic sources and websites to provide accurate answers with line-by-line source citat\u2026 More info can be found at https://knownagents.com/agents/linerbot"
"description":"Manus-User is a browser-enabled AI agent operated by Butterfly Effect, a company based in China. It autonomously navigates websites, interprets content, and carries out m\u2026 More info can be found at https://knownagents.com/agents/manus-user"
"function":"Used to train models and improve products.",
"frequency":"No information.",
"description":"\"The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly.\""
"description":"Meta-ExternalAgent is a web crawler used by Meta to download training data for its AI models and improve its products by indexing content directly. More info can be found at https://knownagents.com/agents/meta-externalagent"
"description":"meta-externalfetcher is used by Meta to perform user-initiated fetches of individual links from AI assistant product functions. More info can be found at https://knownagents.com/agents/meta-externalfetcher"
"description":"Meta-ExternalFetcher is dispatched by Meta AI products in response to user prompts, when they need to fetch an individual links. More info can be found at https://knownagents.com/agents/meta-externalfetcher"
"description":"As per their documentation, \"The Meta-WebIndexer crawler navigates the web to improve Meta AI search result quality for users. In doing so, Meta analyzes online content to enhance the relevance and accuracy of Meta AI. Allowing Meta-WebIndexer in your robots.txt file helps us cite and link to your content in Meta AI's responses.\""
"description":"MistralAI-User is Mistral's AI assistant bot that performs web browsing and data gathering tasks for users in Le Chat, including opening web pages and retrieving informat\u2026 More info can be found at https://knownagents.com/agents/mistralai-user"
"description":"MistralAI-User is for user actions in LeChat. When users ask LeChat a question, it may visit a web page to help answer and include a link to the source in its response.",
"operator":"Naget Inc (founded by Chris Samarinas, headquarter in Amherst, Massachusetts)",
"respect":"Unclear at this time.",
"function":"AI data scraper",
"frequency":"Unclear at this time.",
"description":"'Naget revolutionizes content discovery through an AI-powered ecosystem that transforms how we generate, organize, share, and discover valuable content.' (https://naget.com/) User-agent string links https://naget.ai/bot which yields 404."
"description":"netEstate Imprint Crawler is an AI data scraper operated by netEstate. If you think this is incorrect or can provide additional detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/netestate-imprint-crawler"
"description":"NotebookLM is an AI-powered research and note-taking assistant that helps users synthesize information from their own uploaded sources, such as documents, transcripts, or web content. It can generate summaries, answer questions, and highlight key themes from the materials you provide, acting like a personalized research companion built on Google's Gemini model. NotebookLM fetches source URLs when users add them to their notebooks, enabling the AI to access and analyze those pages for context and insights. More info can be found at https://knownagents.com/agents/google-notebooklm"
"description":"Nova Act is an AI agent created by Amazon that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on behalf of a\u2026 More info can be found at https://knownagents.com/agents/novaact"
"description":"OpenCode is an open-source AI coding agent that helps developers write code from the terminal, IDE, or desktop, supporting multiple LLM providers and local models. More info can be found at https://knownagents.com/agents/opencode"
"description":"Operator is an AI agent created by OpenAI that can use a web browser. It can intelligently navigate and interact with websites to complete multi-step tasks on behalf of a human user. More info can be found at https://knownagents.com/agents/operator"
"description":"PanguBot is a web crawler operated by the Chinese company Huawei. It's used to download training data for its multimodal LLM (Large Language Model) called PanGu. More info can be found at https://knownagents.com/agents/pangubot"
"description":"Perplexity-User supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a \u2026 More info can be found at https://knownagents.com/agents/perplexity-user"
"description":"Phind is an AI-powered answer engine designed for developers, offering technical answers and code examples. It uses real-time web search and specialized AI models to prov\u2026 More info can be found at https://knownagents.com/agents/phindbot"
"description":"Poggio-Citations is a web crawler operated by Poggio, a company that provides AI sales enablement tools for creating tailored narratives, business cases, and account plan\u2026 More info can be found at https://knownagents.com/agents/poggio-citations"
"description":"QualifiedBot is Qualified's web crawler that analyzes customer websites to provide contextual information for their AI-powered chatbots and conversational marketing platf\u2026 More info can be found at https://knownagents.com/agents/qualifiedbot"
"description":"Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg\u2026 More info can be found at https://knownagents.com/agents/querit-searchbot"
"description":"QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time \u2026 More info can be found at https://knownagents.com/agents/queritbot"
"description":"\"AI and machine learning applications often need large amounts of quality data, and web data extraction is a fast, efficient way to build structured data sets.\"",
"description":"Shap-User accesses web content on behalf of users of Parallel Web Systems products. It identifies user-initiated requests rather than automatic web crawling. More info can be found at https://knownagents.com/agents/shap-user"
"description":"ShapBot is a web crawler by Parallel that collects and structures web content to power its search, extraction, and deep research APIs, providing AI agents with high-accur\u2026 More info can be found at https://knownagents.com/agents/shapbot"
"description":"TavilyBot is a web crawler by Tavily that indexes and extracts content from billions of pages, providing real-time search, extraction, and research data to ground AI agen\u2026 More info can be found at https://knownagents.com/agents/tavilybot"
"description":"Terra Cotta is Ceramic's web crawler that indexes public content to power their web-scale search API for AI and LLMs. More info can be found at https://knownagents.com/agents/terra-cotta"
"description":"TerraCotta is Ceramic's web crawler that indexes public content to power their web-scale search API for AI and LLMs. More info can be found at https://knownagents.com/agents/terracotta"
"operator":"Alibaba that fetches web content for the Tongyi Qianwen assistant and related Qwen-generated answers",
"respect":"Unclear at this time.",
"function":"AI Assistants",
"frequency":"Unclear at this time.",
"description":"TongyiBot is a web crawler operated by Alibaba that fetches web content for the Tongyi Qianwen assistant and related Qwen-generated answers. More info can be found at https://knownagents.com/agents/tongyibot"
"description":"Trae is an AI-powered coding agent developed by ByteDance that can understand codebases, fetch web content, and generate code. More info can be found at https://knownagents.com/agents/trae"
"operator":"Twin, a platform that creates automated workers to perform tasks by integrating with APIs and controlling web applications through browser automa\u2026",
"description":"TwinAgent is operated by Twin, a platform that creates automated workers to perform tasks by integrating with APIs and controlling web applications through browser automa\u2026 More info can be found at https://knownagents.com/agents/twinagent"
"description":"UseAI is a web crawler associated with Use AI, a platform that provides an AI workspace where users can chat with AI models, research the web, and perform various tasks. \u2026 More info can be found at https://knownagents.com/agents/useai"
"description":"WARDBot is an AI data scraper operated by WEBSPARK. It's not currently known to be artificially intelligent or AI-related. If you think that's incorrect or can provide more detail about its purpose, please contact us. More info can be found at https://knownagents.com/agents/wardbot"
"description":"Webzio-Extended is a web crawler used by Webz.io to maintain a repository of web crawl data that it sells to other companies, including those using it to train AI models. More info can be found at https://knownagents.com/agents/webzio-extended"
"respect":"Unclear at this time; opt out provided via [Google Form](https://forms.gle/ajBaxygz9jSR8p8G9)",
"function":"Live chat support and lead generation.",
"frequency":"Unclear at this time.",
"description":"wpbot is a used to support the functionality of the AI Chatbot for WordPress plugin. It supports the use of customer models, data collection and customer support."
"function":"According to the [Meltwater Consumer Intelligence page](https://www.meltwater.com/en/suite/consumer-intelligence) 'By applying AI, data science, and market research expertise to a live feed of global data sources, we transform unstructured data into actionable insights allowing better decision-making'.",
"frequency":"Unclear at this time.",
"description":"Retrieves data used for Meltwater's AI enabled consumer intelligence suite"
"operator":"Baidu that fetches web content for the yiyan",
"respect":"Unclear at this time.",
"function":"AI Assistants",
"frequency":"Unclear at this time.",
"description":"YiyanBot is a web crawler operated by Baidu that fetches web content for the yiyan.baidu.com assistant and related ERNIE-generated answers. More info can be found at https://knownagents.com/agents/yiyanbot"