ai.robots.txt/README.md

124 lines
7.4 KiB
Markdown
Raw Normal View History

2024-06-12 19:01:23 -07:00
# ai.robots.txt
2024-03-27 10:58:49 -07:00
<img src="/assets/images/noai-logo.png" width="100" alt="No AI entry logo" />
2024-03-28 09:06:46 +00:00
This list contains AI-related crawlers of all types, regardless of purpose. We encourage you to contribute to and implement this list on your own site. See [information about the listed crawlers](./table-of-bot-metrics.md) and the [FAQ](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/FAQ.md).
2024-03-27 10:58:49 -07:00
2026-07-08 15:30:25 -04:00
A number of these crawlers have been sourced from [Known Agents](https://knownagents.com) and we appreciate the ongoing effort they put in to track these crawlers.
2024-04-08 20:29:45 +01:00
If you'd like to add an AI-related crawler to the list, please see "Contributing" below.
2024-03-29 12:14:48 -07:00
2025-01-17 21:25:23 +01:00
## Usage
This repository provides the following files:
2025-01-17 21:25:23 +01:00
- `robots.txt`
- `.htaccess`
2025-03-27 12:28:11 +01:00
- `nginx-block-ai-bots.conf`
- `Caddyfile`
2025-04-28 08:42:52 +02:00
- `haproxy-block-ai-bots.txt`
2026-02-13 19:20:12 +01:00
- `lighttpd-block-ai-bots.conf`
2025-01-17 21:25:23 +01:00
`robots.txt` implements the Robots Exclusion Protocol ([RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html)).
2025-01-17 21:25:23 +01:00
`.htaccess` may be used to configure web servers such as [Apache httpd](https://httpd.apache.org/) to return an error page when one of the listed AI crawlers sends a request to the web server.
Note that, as stated in the [httpd documentation](https://httpd.apache.org/docs/current/howto/htaccess.html), more performant methods than an `.htaccess` file exist.
2025-01-17 21:25:23 +01:00
2025-03-27 12:28:11 +01:00
`nginx-block-ai-bots.conf` implements a Nginx configuration snippet that can be included in any virtual host `server {}` block via the `include` directive.
`Caddyfile` includes a Header Regex matcher group you can copy or import into your Caddyfile, the rejection can then be handled as followed `abort @aibots`
`haproxy-block-ai-bots.txt` may be used to configure HAProxy to block AI bots. To implement it:
2025-04-28 08:42:52 +02:00
1. Add the file to the config directory of HAProxy
2. Add the following lines in the `frontend` section:
2025-04-28 08:42:52 +02:00
```
acl ai_robot hdr_sub(user-agent) -i -f /etc/haproxy/haproxy-block-ai-bots.txt
http-request deny if ai_robot
```
2025-04-28 09:09:40 +02:00
(Note that the path of the `haproxy-block-ai-bots.txt` may be different in your environment.)
2025-01-17 21:25:23 +01:00
2026-02-15 11:20:34 +01:00
`lighttpd-block-ai-bots.conf` can be included with `include "fragments/lighttpd-block-ai-bots.conf"` in your lighttpd configuration either globally or in any conditional section.
[Bing uses the data it crawls for AI and training, you may opt out by adding a `meta` tag to the `head` of your site.](./docs/additional-steps/bing.md)
2025-05-14 14:11:56 +02:00
### Related
- [Robots.txt Traefik plugin](https://plugins.traefik.io/plugins/681b2f3fba3486128fc34fae/robots-txt-plugin):
middleware plugin for [Traefik](https://traefik.io/traefik/) to automatically add rules of [robots.txt](./robots.txt)
file on-the-fly.
2025-11-26 12:07:39 +00:00
- Alternatively you can [manually configure Traefik](./docs/traefik-manual-setup.md) to centrally serve a static `robots.txt`.
- [Bot Ledger](https://farrelldan.github.io/ai-bot-directory/): free, static directory of verified AI crawlers with a one-click `robots.txt` and `llms.txt` generator. No signup required.
- [KI-Zugangsindex](https://peppe1337.github.io/ki-zugangsindex/): open dataset on how widely this kind of blocking is actually deployed in the German (`.de`) web, measured on a fixed panel of 600 domains so the same sites can be re-checked over time.
2026-09-07 08:09:25 -07:00
- [AI Crawler Census](https://ai-visibility.lastminutedealshq.com/data): open dataset measuring which of these crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run so the same sites can be compared over time. Raw JSON, CC BY 4.0.
- [AI Discovery Radar](https://github.com/flober81/ai-discovery-radar): monthly measurement of how many websites actually publish the files that tell AI systems what they may read or use (`robots.txt`, `llms.txt`, `ai.txt`, `tdmrep.json` and similar), and whether those files can be fetched at all. Open data, CC BY 4.0.
2024-08-02 09:31:48 -07:00
## Contributing
2026-09-19 15:43:27 +01:00
Please note that AI and LLM contributions are not permitted.
2025-03-27 12:28:11 +01:00
A note about contributing: updates should be added/made to `robots.json`. A GitHub action will then generate the updated `robots.txt`, `table-of-bot-metrics.md`, `.htaccess` and `nginx-block-ai-bots.conf`.
2024-08-02 09:31:48 -07:00
2025-11-27 12:41:37 +00:00
You can run the tests by [installing](https://www.python.org/about/gettingstarted/) Python 3, installing the dependencies:
```console
pip install -r requirements.txt
2025-11-27 12:41:37 +00:00
```
2025-11-27 12:41:37 +00:00
and then issuing:
2025-11-27 12:41:37 +00:00
```console
code/tests.py
```
The `.editorconfig` file provides standard editor options for this project. See [EditorConfig](https://editorconfig.org/) for more information.
2025-11-23 04:01:43 +00:00
## Releasing
Admins may ship a new release `v1.n` (where `n` increments the minor version of the current release) as follows:
- Navigate to the [new release page](https://github.com/ai-robots-txt/ai.robots.txt/releases/new) on GitHub.
- Click `Select tag`, choose `Create new tag`, enter `v1.n` in the pop-up, and click `Create`.
- Enter a suitable release title (e.g. `v1.n: adds user-agent1, user-agent2`).
- Click `Generate release notes`.
- Click `Publish release`.
2025-11-23 04:01:43 +00:00
A GitHub action will then add the asset `robots.txt` to the release. That's it.
2024-06-21 22:02:54 -04:00
## Subscribe to updates
You can subscribe to list updates via RSS/Atom with the releases feed:
`https://github.com/ai-robots-txt/ai.robots.txt/releases.atom`.
2024-06-21 22:02:54 -04:00
You can subscribe with [Feedly](https://feedly.com/i/subscription/feed/https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [Inoreader](https://www.inoreader.com/?add_feed=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [The Old Reader](https://theoldreader.com/feeds/subscribe?url=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), [Feedbin](https://feedbin.me/?subscribe=https://github.com/ai-robots-txt/ai.robots.txt/releases.atom), or any other reader app.
Alternatively, you can also subscribe to new releases with your GitHub account by clicking the ⬇️ on "Watch" button at the top of this page, clicking "Custom" and selecting "Releases".
2025-09-14 19:41:36 +08:00
## License content with RSL
It is also possible to license your content to AI companies in `robots.txt` using
the [Really Simple Licensing](https://rslstandard.org) standard, with an option of
2025-09-14 19:41:36 +08:00
collective bargaining. A [plugin](https://github.com/Jameswlepage/rsl-wp) currently
implements RSL as well as payment processing for WordPress sites.
2024-08-10 08:22:28 +08:00
## Report abusive crawlers
If you use [Cloudflare's hard block](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click) alongside this list, you can report abusive crawlers that don't respect `robots.txt` [here](https://docs.google.com/forms/d/e/1FAIpQLScbUZ2vlNSdcsb8LyTeSF7uLzQI96s0BKGoJ6wQ6ocUFNOKEg/viewform).
But even if you don't use Cloudflare's hard block, their list of [verified bots](https://radar.cloudflare.com/traffic/verified-bots) may come in handy.
2024-06-20 08:09:21 -07:00
## Additional resources
- [Blocking Bots with Nginx](https://rknight.me/blog/blocking-bots-with-nginx/) by Robb Knight
- [Blockin' bots.](https://ethanmarcotte.com/wrote/blockin-bots/) by Ethan Marcotte
- [Blocking Bots With 11ty And Apache](https://flamedfury.com/posts/blocking-bots-with-11ty-and-apache/) by fLaMEd fury
2024-06-20 08:11:38 -07:00
- [Blockin' bots on Netlify](https://www.jeremiak.com/blog/block-bots-netlify-edge-functions/) by Jeremia Kimelman
- [Blocking AI web crawlers](https://underlap.org/blocking-ai-web-crawlers) by Glyn Normington
- [Block AI Bots from Crawling Websites Using Robots.txt](https://originality.ai/ai-bot-blocking) by Jonathan Gillham, Originality.AI
- [AI Access Checker: see which AI crawlers a site's robots.txt allows or blocks](https://www.greadme.com/ai-access-checker) by Saar Twito, Greadme