QwenBot, ERNIEBot, DoubaoBot and MistralAI-Index appear in the robots.txt
of Dutch news sites, measured against 83 domains in August 2026, but were
missing from this list.
MistralAI-Index is documented by the vendor at https://docs.mistral.ai/robots
and is the third of Mistral's three crawlers.
For QwenBot, ERNIEBot and DoubaoBot the evidence is different in kind and
the fields say so: the operator is identifiable from the name and the
company's own model family, but none of the three is documented on a
vendor crawler page I could find, so "respect" stays "Unclear at this
time." They are included because publishers are already naming them,
which is the same reason this list carries other undocumented agents.
Only robots.json is changed; the other files are regenerated
automatically.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
All three are documented by the vendor itself:
- KimiBot and Kimi-SearchBot: https://www.kimi.ai/policies/kimi-crawlers
The list already has Kimi-User; Moonshot documents three crawlers.
- MistralAI-Training: https://docs.mistral.ai/robots/
The list already has MistralAI-User. Mistral documents three; the
third, MistralAI-Index, is proposed separately in #274.
Only robots.json is changed, per review: the other files are regenerated
automatically. Rebased on main so that Diffbot-User, merged in the
meantime, is preserved.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
Added at a maintainer's request in #277, where I raised it but did not
propose it.
OpenAI documents it at https://developers.openai.com/api/docs/bots as the
fourth of its four crawlers. Two things in that entry are worth recording
in the fields rather than leaving to the reader:
- The page does not say OAI-AdsBot obeys robots.txt, and gives no
disallow example, so "respect" is "Unclear at this time." rather than
the "Yes" the other three OpenAI crawlers carry.
- OpenAI states the data is not used to train foundation models, and that
the crawler only visits pages submitted as ads. What makes it belong on
this list is the other half of the same paragraph: content from the
landing page is also used to "determine when it's most relevant to show
the ad to users".
Only robots.json is changed; the other files are regenerated
automatically.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141w8LsSMs7AHbon2ySf84n
Diffbot documents two separate user agents. The list currently has only
the proactive crawler; this adds the user-triggered one.
From Diffbot's own documentation:
Diffbot-User — This is used by requests originating on behalf of a
human user browsing a URL using Diffbot software, in response to
their input.
Source: https://docs.diffbot.com/docs/does-crawl-respect-robotstxt
Same distinction as GPTBot / ChatGPT-User and PerplexityBot /
Perplexity-User, so it is classified as an AI Assistant fetching only
when prompted by a user.
Observed in the robots.txt of five independent publishers in four
countries: nytimes.com, bbc.co.uk, lefigaro.fr, corriere.it and
usatoday.com.
Generated files updated with code/robots.py --convert.
The Lightpanda browser is an AI-centric headless browser used for scraping and automation.
It is currently being used by a botnet of millions of IP addresses, including over 20k SpaceX IPs, to scrape sites such as the official Haskell Gitlab:
https://github.com/lightpanda-io/browser/issues/3156
Re-submission of #248, which was merged then reverted in 80c19fc. The revert was
caused by main.yml's commit-message handling, not by this data — that is fixed
separately in #254.
Six entries move from "Unclear at this time." to the operator's own documented
value, each citing the primary source inline:
Applebot Yes support.apple.com/en-us/119829#retrieval
Meta-ExternalAgent Yes developers.facebook.com/docs/sharing/webmasters/web-crawlers/
Meta-ExternalFetcher No (same Meta page — user-initiated fetch, documented
meta-externalfetcher No as not checking robots.txt)
DuckAssistBot Yes duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/
Google-Agent Yes developers.google.com/search/docs/crawling-indexing/google-common-crawlers#google-agent
table-of-bot-metrics.md is regenerated and included, so `robots.py --convert` has
nothing left to write and the workflow's commit step is skipped entirely.
Verified: code/tests.py 13/13 · robots.py --convert exits 0 and is idempotent ·
robots.txt, .htaccess, the nginx/lighttpd/haproxy configs and the Caddyfile are
byte-identical · the JSON diff touches only the six `respect` fields, nothing else.
Applebot, Meta-ExternalAgent, Meta-ExternalFetcher, meta-externalfetcher,
DuckAssistBot and Google-Agent carried 'Unclear at this time.' in both
operator and respect, but each operator documents the behaviour on its
own site. Values are linked to the primary source inline, matching the
style already used by Applebot-Extended and meta-externalagent.
robots.txt output is unchanged — no keys added or renamed.
User-agent header values seen in the wild over the last days:
Mozilla/5.0 (compatible; NagetBot/1.0; +https://naget.ai/bot)
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 newsai/1.0 Safari/537.36