Every guide to blocking AI crawlers comes with a list. Twenty names, thirty names, all in capital letters, all apparently reading your site right now. Which of them actually turn up on an ordinary WordPress site, how often, and which ones are worth a decision?
We counted. Our plugin recognises 26 names, and it recorded every visit on two WordPress sites we run, a news site for nine days and a travel guide for 30. This article ranks what came, names what never did, counts the visits that used a real bot’s name from the wrong address, and sends you into the bot directory for any name you want to know more about.
The short answer
- Of the 24 real bots on the list (two of the 26 names are opt-outs, not bots), 11 visited the news site in nine days and 10 visited the travel site in 30. Thirteen never reached either site.
- Four bots did most of the visiting. Applebot, Amazonbot, ChatGPT-User and Claude-User were nine visits in ten on the news site and seven in ten on the travel site.
- On the busier site, one visit in 65 used a bot’s name from an address its company does not publish. All of those used the two live-read names.
- The mix depends on the site. A news site is read live by assistants all day. A travel guide is harvested by training crawlers and then left alone.
You do not need 26 rules. You need to know which job each name does, and then to look at your own log. The rest of this article shows what ours said.
The 26 AI crawlers, sorted by job
Every name on the list does one of three jobs, and the job matters more than the company. Live-read bots fetch a page because a person just asked an assistant a question. Search crawlers build an index so an assistant can find and link your pages later. Training crawlers copy pages to build future models.
Two more names on the list are not bots at all but opt-outs, and they are explained at the end. Each name below links to its page in the directory.
Live read. ChatGPT-User (OpenAI), Claude-User (Anthropic), Perplexity-User (Perplexity), meta-externalfetcher (Meta), MistralAI-User (Mistral AI), DuckAssistBot (DuckDuckGo), Grok-DeepSearch (xAI).
Search crawl. OAI-SearchBot (OpenAI), Claude-SearchBot (Anthropic), PerplexityBot (Perplexity), Applebot (Apple), YouBot (You.com), xAI-Grok (xAI).
Training crawl. GPTBot (OpenAI), ClaudeBot (Anthropic), GoogleOther (Google), CCBot (Common Crawl), Bytespider (ByteDance), Amazonbot (Amazon), meta-externalagent (Meta), cohere-ai (Cohere), Diffbot (Diffbot), Timpibot (Timpi), GrokBot (xAI).
Two of those need a word of caution. Applebot is Apple’s ordinary search crawler, the one behind Siri, Spotlight and Safari suggestions, and it has been around far longer than the AI wave. It is on the list because Apple may also use what it collects to train its models unless you opt out with the Applebot-Extended name.
Amazonbot is Amazon’s general crawler, which Amazon says improves its products and may be used to train its AI models. When either of them tops your ranking, most of that is plain search crawling.
What actually turned up, ranked
Here is the whole list of AI crawlers against both sites. The news site column covers 16 to 24 September 2026, nine days. The travel site column covers 30 days to 24 September. The windows differ, so read down each column rather than across, and remember both sites sit behind a full-page cache, a saved copy of each page served without running WordPress, which hides some bot visits from the plugin altogether.

| Bot and job | News site, 9 days | Travel site, 30 days |
|---|---|---|
| Applebot, search | 2,510 | 394 |
| ChatGPT-User, live read | 2,172 | 124 |
| Claude-User, live read | 1,858 | 92 |
| Amazonbot, training | 1,361 | 754 |
| OAI-SearchBot, search | 353 | 78 |
| PerplexityBot, search | 193 | 295 |
| YouBot, search | 124 | 76 |
| ClaudeBot, training | 121 | 82 |
| GoogleOther, training | 86 | 14 |
| DuckAssistBot, live read | 47 | 32 |
| GPTBot, training | 27 | 0 |
| The other 13 bots | 0 | 0 |
| All AI bot visits | 8,852 | 1,941 |
The 13 that never reached WordPress on either site were Claude-SearchBot, Perplexity-User, CCBot, Bytespider, meta-externalagent, meta-externalfetcher, cohere-ai, Diffbot, Timpibot, MistralAI-User, GrokBot, xAI-Grok and Grok-DeepSearch. GPTBot managed 27 visits to the news site and none to the travel site in a month.
A news site is read live all day
The news site publishes many new pages a day, and the assistants respond to that. Live reads were the largest group, 4,077 of the 8,852 visits, or 46 percent. ChatGPT-User and Claude-User between them fetched pages more than 4,000 times in nine days, which is people asking questions and assistants going to look. Search crawls were 3,180 and training crawls 1,595. The two big training names, GPTBot and ClaudeBot, made 148 visits between them, less than two percent.
Applebot’s 2,510 visits touched 1,502 different pages, so it was working through the whole site rather than re-reading the same few. ChatGPT-User’s 2,172 visits touched 798 pages, and Claude-User’s 1,858 touched 630, which is what you would expect when the same fresh stories get asked about repeatedly.
A travel guide is harvested and then left alone
The travel site changes slowly, and the mix flips. Training crawls were the largest group, 850 of 1,941 visits, and Amazonbot alone was 754 of those, 39 percent of everything. Search crawls were 843, with PerplexityBot’s 295 outranking every live-read bot. Live reads were 248, one visit in eight. A page that does not change gets read once for training and then only when someone asks about it.
Per day, the difference is stark. The news site averaged about 980 AI bot visits a day, the travel site about 65. Two sites cannot tell you what a shop or a brochure site will see, and ours are both cached, so treat these as two worked examples rather than a benchmark.
These counts are floors. Both sites serve most pages from a cache, and a bot visit answered by the cache never reaches WordPress, so the plugin never sees it. A bot that only reads popular, cached pages could be under-counted badly. Our caching guide explains how to recover those visits from the server log.
Why your site will not look like the headline numbers
Network-wide figures tell a different story, and the gap is instructive. TechnologyChecker’s analysis of Cloudflare Radar data for July 2026, the latest full month when we checked, ranks AI-related crawlers by share of that traffic. Googlebot leads at 24.5 percent, then ClaudeBot 16.3, meta-externalagent 12.7, GPTBot 9.7, Bingbot 8.9, Applebot 6.5, Amazonbot 6.0, Bytespider 5.1 and Claude-SearchBot 3.5.
Googlebot and Bingbot appear because the same crawl feeds Google’s and Microsoft’s AI features. Our plugin leaves them out on purpose, because blocking them removes you from search.
Set that against our two sites. ClaudeBot, second across the network, was eighth on the news site. meta-externalagent and Bytespider, third and eighth across the network, never reached WordPress on either site. Applebot and Amazonbot, sixth and seventh across the network, were first and fourth here. And the live-read bots that filled our logs barely feature in crawler rankings at all, because a crawl of a million pages counts for more than a thousand single fetches.
Where the difference comes from
Three things explain most of it. Network data counts every request at the edge, the network layer in front of your server, including the ones a cache answers and the ones a firewall refuses, while our plugin sees only what reaches WordPress.
Training crawlers spend their effort on sites with enormous numbers of pages, so a small site gets a short visit and silence. And a news site is asked about far more than it is crawled, which pushes the live-read bots to the top.
Bytespider and meta-externalagent deserve one more sentence. Some hosts and content delivery networks, the services that serve cached copies of your pages, block Bytespider before it reaches WordPress, and both are bulk crawlers whose requests a full-page cache is likely to answer. Never reaching WordPress is not the same as never visiting. The only way to be sure is to import the server log, which sees everything.
One visit in 65 used a name from an address its company does not list
A bot’s name is just a string it sends. Anyone can send it. Where a company publishes its bots’ addresses, or lets them be confirmed by looking up the address’s registered name, the plugin checks every visit and labels it Verified, Spoofed when the check fails, or Unverified when the list could not be fetched and no recent saved copy existed. For companies that publish nothing, it shows No check rather than guess.
On the news site, 135 of 8,852 visits were Spoofed, 1.5 percent. Every one of them used a live-read name. 123 called themselves Claude-User and 12 ChatGPT-User. Not a single visit faked a crawler name. On the travel site it was 7 of 1,941, five ChatGPT-User and two Claude-User.
Spoofed does not mean attacker. The 135 visits came from dozens of unrelated addresses, one to five visits each, many of them in New Zealand home broadband ranges. That pattern does not look like one impostor.
One reading is that people’s own tools are sending the Claude-User name from home connections, which a list of Anthropic’s server addresses would not cover. A scanner, a scraper hiding behind a trusted name, or a new address the company has not listed yet are the other possibilities. The label tells you the address did not match. What to make of that is a judgement, and our ClaudeBot article goes into it.
What this means for your robots.txt
Neither site blocks anything yet, and the numbers say why the decision is smaller than the lists make it look. The 13 names that never came cost nothing to block in advance and change nothing if you do.
Four of the 24, Bytespider, Perplexity-User, xAI-Grok and Grok-DeepSearch, are recorded as not always obeying robots.txt, so a line for them is a request, not a lock. The real decision is the handful of training crawlers that actually visit, which for these sites means Amazonbot, ClaudeBot, GoogleOther and GPTBot.
Two cautions before you flip switches. Blocking Applebot or Amazonbot as if they were AI bots also removes you from Siri and Spotlight, and from whatever Amazon builds on its crawl, because those are their day jobs.
The two opt-out names, Google-Extended and Applebot-Extended, never appear in a log, because they are not bots. They are lines you can add to decline training by Google and Apple while their ordinary crawlers carry on. Our guide to whether to block AI bots sets out the rule we use, which is to decide by job, not by list.
See your own list
Everything above came from the free Forge AI Bot Log plugin’s dashboard and activity log, plus the pages-touched counts from its Pro add-on. Install it, wait a month, and your ranking will be different from ours, which is the point. Then use the directory to read up on the names that actually appear, and ignore the rest until they do.
Common questions
Is Googlebot an AI crawler?
Google uses the same crawl for Search, AI Overviews and AI Mode, so network data often files it under AI. Blocking it removes you from Google Search. The plugin does not count it. If you want to decline training use of what Google already crawls, the Google-Extended name does that without touching Search.
Why does Applebot count as an AI bot?
Because Apple may use its crawl to train its models unless you add the Applebot-Extended opt-out. Its main job is still search for Siri, Spotlight and Safari, which is why it tops both our sites.
Should I block the bots that never visit?
You can, and it costs nothing today. It also achieves nothing today. Add the lines if you want a tidy policy, and spend your attention on the training crawlers that actually show up in your log.
Count the bots on your own site
The free plugin records every visit from the 26 AI crawlers, labels each by job, checks the address where the company allows it, and gives you one switch per bot. Pro adds the people those assistants send you.