In 2026, AI traffic is counted on a declaration: a bot is whatever it says it is. This Fast Growth Advisors white paper measures verified AI traffic by applying to thirteen days of server logs the identity verification that publishers make possible. Of 1,129 requests presenting themselves as an AI assistant, 876 were proven impersonations. The central result fits in one line: the error rate blamed on AI crawlers mostly belongs to those impersonating them.
Fast Growth Advisors · August 2026
What is verified AI traffic, in brief?
It is the share of AI bot visits whose identity has been confirmed, rather than merely announced by the bot itself.
Yet the AI visibility measurement market rests on a convention nobody states out loud: a bot is whatever it declares itself to be. Every automated visitor announces its identity in a line of text it writes itself, the user agent. Almost every published figure on AI crawling, scrape ratios, error rates, crawler market shares, adds up those declarations.
At Fast Growth Advisors, we applied to our own server logs, for thirteen days, the identity verification that engine publishers make possible: published address ranges, reverse resolution confirmed in both directions.
No ambiguity in the verdict.
Of 1,129 requests presenting themselves as an AI assistant fetching a page, 876 were proven impersonations, and they are what produces most of the error rate a classic tool would report. Authenticated requests, by contrast, almost never knock on the wrong door: zero missing pages across 154 requests.
This document describes the method, gives every figure with its base and states the limits of our own instrument. It ends with the one question that lets anyone assess an AI visibility tool in a single sentence.
Is a declaration an identity?
No. Dating back to the early web, the user agent is a free text field: nothing stops any machine from writing Googlebot or ChatGPT-User in it.
Many do, because the disguise opens doors. A firewall lets through what looks like a search engine, and a site accepts being scraped by what it believes to be an index.
System administrators have known the consequence for twenty years; almost every marketing dashboard ignores it: counting user agents means counting declarations. The resulting figure mixes genuine bots, impersonators, and anything disguising itself to get through.
There is a way to settle it, bot by bot, request by request, and the Fast Growth Advisors instrument relies on three mechanisms, from strongest to weakest. On one side, cryptographic signature, defined by Web Bot Auth and RFC (Request For Comments) 9421, still barely deployed. Next, the address range published by the vendor, which OpenAI makes available as JSON (JavaScript Object Notation) at stable addresses. And reverse resolution confirmed in both directions: you ask for the name attached to the address, then check that the name points back to that address. One way is not enough, an impersonator can configure the return. Only the round trip holds.
Two implementation traps, learned by doing. The expected domain suffix must start with a dot, otherwise evil-googlebot.com satisfies a googlebot.com test. As for real suffixes, they are measured, not assumed: Amazon’s bot resolves to .crawl.amazonbot.amazon, whose top level domain is .amazon and not .amazon.com.
Whoever assumes lets them through.
Why four verdicts rather than two?
Because the binary reflex, verified or not, mixes two morally opposite situations: the bot that fails a possible verification, and the bot that no mechanism allows you to verify at all.
Our instrument at Fast Growth Advisors therefore assigns four verdicts to every request. Authenticated: identity confirmed by one of the three mechanisms. Proven impersonation: verification was possible and it failed, the address does not belong to the declared vendor. Unverifiable: the vendor publishes neither ranges nor reverse records, no test is possible. Undecided: verification should have succeeded and did not, which flags a limit of the instrument rather than a fault of the bot.
A cosmetic distinction? No.
ClaudeBot, Anthropic’s crawler, is unverifiable: 841 requests on our main site during the window, none verified and no proven impersonation either. It is not an impersonator, it is a bot without verifiable papers. The same holds for meta-externalagent, youbot and claude-user, and for the DuckDuckGo bot, at zero verification on both sites.
It also separates names one might have thought neighbours. bytespider and google-extended are not unverifiable: their verification does resolve, and it fails. 133 requests out of 133 for the first, 292 out of 292 for the second, all proven false. Carrying a known vendor’s name therefore says nothing about what sits behind it.
This list does not ask who cheats. It asks why some vendors publish what is needed to verify their bots while others publish nothing at all.
Publishing costs nothing, and the benefit, for them as for the sites they visit, is immediate.
What does verification show over thirteen days?
Our measurements cover the window from 11 to 23 August 2026, thirteen full days, on the server logs of two B2B sites, fast-growth.fr and nomo-ia.com.
The window was chosen clean: after our classification fixes of 8 August and the reclassification of history, before an infrastructure change on 24 August, with no deployment in between. Detail and aggregates agree across the whole window, with no discrepancy.
First Fast Growth Advisors table: declared against verified, by bot class.
| Declared class | fast-growth.fr | Verified | nomo-ia.com | Verified |
|---|---|---|---|---|
| Classic search engine | 492 | 78.9% | 567 | 96.5% |
| AI index bot | 1,419 | 51.1% | 882 | 46.4% |
| Training collection | 3,907 | 9.5% | 3,198 | 13.8% |
| User-triggered fetch | 1,129 | 13.6% | 688 | 3.5% |
Overall, the reading is clear: classic engines are verifiable at 79% and 97%, AI agents at 3% and 14%. Same method, same window, same server.
But that global rate blends the four verdicts, and the breakdown says something other than "AI bots are opaque".
| Verdict | fast-growth.fr (6,455 req.) | Share | nomo-ia.com (4,768 req.) | Share |
|---|---|---|---|---|
| Authenticated | 1,250 | 19.4% | 875 | 18.4% |
| Proven impersonation | 3,372 | 52.2% | 3,646 | 76.5% |
| Unverifiable, silent vendor | 1,041 | 16.1% | 239 | 5.0% |
| Undecided, instrument limit | 792 | 12.3% | 8 | 0.2% |
In the middle sits the figure that matters. Most declared AI crawling is not unverifiable, it is proven false: 52% on one site, 76% on the other. As for the silent vendor nuance, the one honesty requires, it covers only 16% and 5% of the volume.
Which bot names serve as a disguise?
On fast-growth.fr, the most impersonated name in the window does not belong to an answer engine: it is amazonbot, with 1,165 proven impersonations across 1,388 requests.
Here is the detail, bot by bot.
| Declared bot | Vendor | Requests | Authenticated | Impersonations | Impersonated share | Dominant verdict |
|---|---|---|---|---|---|---|
| amazonbot | Amazon | 1,388 | 223 | 1,165 | 84 % | mixed |
| chatgpt-user | OpenAI | 1,030 | 154 | 876 | 85 % | mixed |
| claudebot | Anthropic | 841 | 0 | 0 | 0 % | unverifiable |
| meta-externalagent | Meta | 792 | 0 | 0 | 0 % | unverifiable |
| gptbot | OpenAI | 452 | 148 | 304 | 67 % | mixed |
| oai-searchbot | OpenAI | 441 | 155 | 286 | 65 % | mixed |
| petalbot | Huawei | 419 | 419 | 0 | 0 % | authenticated |
| perplexitybot | Perplexity | 360 | 52 | 308 | 86 % | mixed |
| google-extended | 292 | 0 | 292 | 100 % | impersonated | |
| googlebot | 231 | 149 | 23 | 10 % | mixed | |
| bingbot | Microsoft | 213 | 206 | 7 | 3 % | mixed |
| bytespider | ByteDance | 133 | 0 | 133 | 100 % | impersonated |
| applebot | Apple | 107 | 99 | 8 | 7 % | mixed |
| claude-user | Anthropic | 99 | 0 | 0 | 0 % | unverifiable |
| youbot | You.com | 92 | 0 | 0 | 0 % | unverifiable |
Three families emerge from this Fast Growth Advisors breakdown. On one side, bots whose identity can be verified and holds: petalbot at 100% authentication, bingbot, applebot, googlebot. Then those no mechanism allows you to verify, therefore never accused: claudebot, meta-externalagent, claude-user, youbot. And those whose name is massively used as a disguise.
Another result deserves flagging: google-extended, the agent Google associates with its training uses, shows no authenticated request at all across 292. On our server and in this window, every request carrying that name is proven false.
With googlebot, from the same vendor, the contrast is instructive: 149 authenticated requests out of 231, and 23 impersonations.
Google lets you verify that bot, and it is verified; the one whose verification is not documented the same way serves as a mask.
Who does the AI crawler error rate belong to?
Mostly to impersonators, at least on our servers.
Fast Growth Advisors isolated the most commented category on the market: the user fetch, meaning a request from an assistant retrieving a page on a human’s behalf. Same window, same base, on fast-growth.fr.
| Verdict | Requests | Pages found | Missing pages |
|---|---|---|---|
| Authenticated | 154 | 99.4% | 0.0% |
| Not verified | 99 | 87.9% | 1.0% |
| Proven impersonation | 876 | 30.7% | 69.1% |
| All verdicts combined | 1,129 | n.a. | 53.7% |
Authenticated, requests find their page. Always, or nearly: not one missing page across 154 requests. Impersonations, by contrast, grope around: 69.1% of their requests target pages that do not exist, the typical behaviour of vulnerability scanning, not of content reading.
And the total, the only reading a non-verifying tool can produce, shows a 53.7% error rate. A figure manufactured entirely by impersonators.
Then comes the external comparison. Vercel, in the December 2024 reference study, publishes 34.82% missing pages for ChatGPT and 34.16% for Claude, against 8.22% for Googlebot. Those rates, quoted everywhere as proof that AI crawlers waste resources, sit in the range of our unfiltered total, never in the range of our authenticated traffic.
So we offer a hypothesis, given for what it is: the error rate attributed to AI crawlers may measure, to a substantial extent, the behaviour of those impersonating them. Of Vercel’s method, we do not know the detail, and we claim nothing about their figures. What we show: on our servers, with a public and replayable method, the gap between the declarative reading and the verified reading reaches 53.7 points.
Why do two sites give different rates?
Impersonators do not spread evenly across the two sites.
In the first table, classic engines are verified at 78.9% on one site and 96.5% on the other. Eighteen points apart, same method, same resolver. An anomaly?
Fast Growth Advisors therefore resolved the addresses one by one. 67 distinct addresses declared themselves Googlebot on fast-growth.fr: 7 with a genuine Google resolution, 3 pointing to third parties, 57 with no resolution at all. On nomo-ia.com, 20 declared addresses, including the same 7 genuine ones.
Identical on both sides, the real Googlebots do not make the difference: the fake ones do.
Among the impersonators, machines hosted with providers where anyone can rent a machine by the hour. The gap is therefore not a flaw in the method. It measures each site’s exposure to impersonators, a quantity no declarative tool can produce, since it does not know they are impersonators. A site that is more visible, older or more attacked carries more fake Googlebots. The control figure does control something.
Do acquisition channels lie too?
Partly, yes. Identity verification corrects the counting of bots, but two blind spots affect every tool on the market, the Fast Growth Advisors one included.
First blind spot, the referral channel. Over the same window, 12,087 requests in our logs carried a referring site, the classic definition of traffic coming from elsewhere. 10,740 of them, 88.9%, were internal navigation: a visitor already on the site clicking through to another page of the same site, which the classification chain filed under acquisition for want of a rule. The fourth channel of our acquisition screen was wrong nearly 90% of the time. We fixed the rule. How many acquisition screens elsewhere still count their own navigation as referred visits?
Second blind spot, AI referrals. Over the window, only four requests were referred by an AI interface. That figure is a floor, not a measurement: the native ChatGPT and Claude applications send no provenance header, so an unknowable share of that traffic arrives as direct. Any "zero AI referrals", ours included, reads as "zero referrals carrying a header", never as "zero visitors sent".
On that note, the scrape ratios in circulation.
| Source | Ratio | Window | Scope |
|---|---|---|---|
| Cloudflare | 70,900 : 1 | June 2025 | Anthropic bots |
| Cloudflare | ≈ 4,580 : 1 | 28 days, late June 2026 | Cloudflare network |
| Microsoft Clarity | ≈ 6,000 : 1 | August 2026 | published example |
Windows, scopes and methods differ. Quoting a ratio without its window is quoting a headline, not a measurement. And since these ratios are computed on undercounted referrals, they mechanically overstate, as their own authors point out.
Who actually gets through, once papers are checked?
OpenAI, Huawei and Amazon, each in its own category.
Once traffic is filtered on confirmed identities, the real landscape is surprising, and the orders of magnitude Fast Growth Advisors measured over seventeen days and two sites, before the main window, have been stable since.
For fetches triggered by a real prompt, OpenAI accounts for all verified traffic. In AI indexing, the leading bot is neither OpenAI nor Anthropic: it is PetalBot, Huawei’s crawler, at 55% and 68% depending on the site. Within training collection, amazonbot runs about forty times the volume of GPTBot.
What matters most remains invisible by construction.
Gemini never visits your server: it reads Google’s index, built by Googlebot for classic search. A fetch dashboard, however well verified, stays structurally blind to one of the market’s main answer engines.
Your tools show you who comes by. They cannot show you who reads you without coming by.
What are the limits of our own measurement?
Four limits, all declared here.
A document that faults the market for its declarative habits owes it to itself to be beyond reproach about its own gaps. At Fast Growth Advisors, the 792 undecided requests on fast-growth.fr, 12.3% of AI agent volume, concern a single bot, meta-externalagent, whose verification succeeds on one site and not the other. Neither an impersonation nor an impossibility, it is a gap in our instrument, unresolved as of August 2026, counted separately and declared here rather than quietly fixed.
Second limit, volumes: those of two small company sites, meaning 154 authenticated fetches in thirteen days on the main site, 24 on the second.
Too few to carry a figure on their own. We are not describing the web: we are describing a method and its effect on the numbers.
Third limit, the window: thirteen days, a short span that will be extended. Last limit, the classification of the categories themselves, index, training, fetch, relies on vendors’ declarations about the role of their bots, declarations we cannot verify.
What question should you ask your measurement tool?
Just one, to be put to a tool or a provider: what does bot identification rest on in this report? For an executive, this whole document folds into that question.
If the answer is the user agent, the volumes shown mix bots and impersonators in proportions this document lets you imagine.
Meanwhile, the standard is moving. Cryptographic signatures from Web Bot Auth, described by RFC 9421, would give strong verification without reverse resolution, and OpenAI already publishes its address ranges at stable locations. The day signing becomes the norm, the question will change sides: it will be put to the bots that do not sign, and to the vendors that publish nothing.
Until then, verification in both directions remains the best tool available, and it is within reach of anyone who reads their server logs. For Fast Growth Advisors, counting correctly is the precondition for any AI visibility strategy: before optimising for bots, make sure they are bots.
What they then read of your pages is the subject of another white paper, the recoverable message.
Frequently asked questions
How do I verify the identity of a bot visiting my site?
Through three mechanisms, from strongest to weakest.
They are the cryptographic signature defined by RFC 9421, the address range published by the vendor, and reverse resolution confirmed in both directions. Reverse resolution means asking for the name attached to the address, then checking that the name points back to that address. One way is not enough: an impersonator can configure the return.
Why is the AI crawler error rate so high in published studies?
Because it is computed on declarations.
In the logs analysed by Fast Growth Advisors, authenticated requests show zero missing pages across 154 requests, against 69.1% for proven impersonations. Without identity verification, only the unfiltered total can be read, and it comes out at 53.7%. A tool that does not verify therefore attributes to engines the behaviour of those imitating them.
Is ClaudeBot suspicious since it is never verified?
No: ClaudeBot is unverifiable, which differs from an impersonation.
Anthropic publishes neither address ranges nor reverse records that would allow the test, so no verdict of any kind can be reached on a single ClaudeBot request. Confusing the two would mean accusing a bot for want of being able to identify it.
Can the scrape ratios published by different players be compared?
Not unless you also compare their windows and scopes.
Three ratios are in circulation, 70,900 to 1, about 4,580 to 1 and about 6,000 to 1. All are real, but they cover different periods, networks and methods. A ratio quoted without its window is a headline, not a measurement.
Is my AI visibility tool reliable?
Ask it what bot identification rests on in its report.
If the answer is the user agent, the volumes mix genuine bots and impersonators. If the answer mentions published address ranges or reverse resolution in both directions, the tool knows what it is talking about.
Glossary
Fast Growth Advisors uses these seven terms in a precise sense, the one of this white paper.
- User agent
Line of text that an automated visitor writes itself to announce its identity. Free field, therefore forgeable: counting it means counting declarations.
- Reverse resolution in both directions
Check that asks for the name attached to an address, then confirms that the name points back to that same address. Only the round trip proves anything, since one way can be configured by an impersonator.
- Proven impersonation
The verdict when verification was possible and failed: the address does not belong to the declared vendor.
- Unverifiable
The verdict when the vendor publishes neither address ranges nor reverse records, to be distinguished from a proven impersonation. No test is possible, no fault is established.
- User fetch
Request from an assistant retrieving a page on a human’s behalf, as opposed to indexing and training collection.
- Web Bot Auth
Mechanism for cryptographically signing bot requests, built on RFC 9421, which would provide strong verification without relying on reverse resolution or on published address ranges.
- Scrape ratio
The ratio between pages scraped by a vendor’s bots and visitors it refers back. Sensitive to window and scope, therefore not comparable across sources.
Sources
Six references in all. For every external figure, Fast Growth Advisors points to the publication of the vendor or standards body that produced it, and lists its own measurements last, with their bases.
- Vercel and MERJ (web analytics firm). The rise of the AI crawler, 17 December 2024. Missing page rates of 34.82% for ChatGPT, 34.16% for Claude and 8.22% for Googlebot. vercel.com
- OpenAI, published address ranges for its bots, in JSON, at stable addresses. gptbot.json, searchbot.json, chatgpt-user.json
- RFC 9421, HTTP Message Signatures, IETF (Internet Engineering Task Force), February 2024. Technical basis for Web Bot Auth. rfc-editor.org
- Cloudflare. From Googlebot to GPTBot: who’s crawling your site in 2025. Ratios between pages scraped and visitors referred. blog.cloudflare.com
- Microsoft Clarity. Scrape-to-Referral insights, 13 August 2026. A ratio of about 6,000 to 1 in the published example. clarity.microsoft.com
- Fast Growth Advisors measurements, server logs of fast-growth.fr and nomo-ia.com, window from 11 to 23 August 2026. Bases stated with every figure in this document.
This document is published in open access and may be cited with attribution.