IA & Marketing

Your site isn’t invisible to AI. It’s unreadable.

A French Series-B deeptech raised 30 million euros. Its process cuts production costs by 45% and carbon footprint by 60%. Those figures are in the trade press and in the funding announcements. None of those figures exists on its own site in a form an engine can reuse. That gap is what AI engine readability measures, and what Fast Growth Advisors checks in every audit.

On its homepage, a visitor reads: « innovative advanced recycling solutions. »

Ask an AI assistant about it.

It will describe what the company does. It won’t say what sets it apart, so it will name the category instead, and cite as a reference the competitor who put its proof in HTML.

This case comes from our audits, anonymized, and there is nothing exceptional about it.

We found the same pattern in 61% of the 83 detailed audits in the Fast Growth Advisors (FGA) Observatory, and our measurements confirm it: the figures that prove the value live in the press and the funding announcements, almost never on the website.

The debate opening up is the wrong one

Since 19 August, Microsoft Clarity has shown a scrape-to-referral ratio: how many pages AI systems extract for every visitor they send back. About 6,000 to 1, on their data. Cloudflare had published 70,900 to 1 for June 2025, then around 4,580 to 1 a year later.

Three real figures, three measurement windows, none comparable. The debate will still crystallize around them: AI takes a lot and returns little, what does it owe us?

A fair question.

But for a B2B startup leader, it misses the only thing that matters: when an engine reads your homepage, what does it keep?

The question is not rhetorical. It can be measured. We measured it.

Three ways to disappear

The most common mistake is to think in one step: « the bot sees the page. » There are three, and a piece of information can survive the first two and die at the third.

First the fetch: the agent requests the URL, and can fail on a firewall, a missing page, a consent banner. Then the response: the server returns bytes, without executing a single line of JavaScript. Finally the extraction: a tool turns that HTML into usable text, and throws away part of the document along the way.

Three stages, three causes of invisibility, three different fixes. An audit that does not separate them produces a false diagnosis.

The third stage is almost always forgotten. Yet that is where the least intuitive results are found.

What the extractors throw away

We built a test page with twenty unique markers, one per edge case, embedded in a body of realistic length. Then we ran it through the three extractors used in the corpus-preparation pipelines of the large models: jusText, Resiliparse and Trafilatura.

Three results are worth pausing on.

A pure client-rendered page, meaning a React application that builds all its content in the browser, serves 128 bytes of HTML. Usable text after extraction: zero bytes. The case is binary, with no partial degradation.

A block hidden with CSS but rendered by the server is read in full. The reason is mechanical: with no JavaScript execution, no CSS is applied, so the engine has no idea the block is hidden. It is the exact inverse of classic SEO, where Google renders the page and can devalue hidden content.

Twenty years of reflexes to unlearn.

And the test everyone recommends is wrong. A curl followed by a grep on the raw response finds your sentence even when it is injected by JavaScript, because it sits as a literal string inside the script, precisely the part every extractor discards. The test clears the very sites it was meant to catch.

The proof leaves first

On a typical startup, text coverage is fairly good. The message gets through. What disappears are the elements that prove: client logos in a JavaScript carousel, reviews in a third-party widget, a monthly-annual pricing toggle, key figures animated on scroll.

The engine can then say what the company does. It cannot say what sets it apart.

Add the 61% pattern noted above, and the mechanism closes: the proof that existed was already somewhere other than the site, and the proof that remained on the site is carried by JavaScript the engines do not read. Invisible twice over.

What do the extractors keep from a real page?

We built a fictional B2B startup homepage carrying the most common defects: result figures, client logos, testimonial and pricing grid injected by JavaScript. Then we ran the three extractors used to prepare the corpora of large models on it.

What three extractors keep from the same page. Measured on 27 August 2026.
Reading Bytes of text Share Tracked elements kept
Seen by a human, JavaScript executed 1,450 100% 8 / 8
trafilatura 706 49% 1 / 8
jusText 796 55% 2 / 8
Resiliparse 949 65% 2 / 8

Volume reassures, detail damns. Half the text survives, but the elements kept are never the ones that matter: the promise gets through, and so does the old positioning left in display:none during a redesign, in all three. The six real proofs, the three result figures, the client logos, the testimonial and the prices, survive nowhere.

The screenshot, the three extractions reproduced in full and the complete protocol are in the white paper: The recoverable message, freely available, with its PDF version.

0.91 out of 5

The FGA Observatory scores fifteen sub-criteria across 94 full audits of post-funding French startups. The weakest sub-criterion of all: readability by AI engines, at 0.91 out of 5. Less than 20% of the potential.

For context, none of the fifteen sub-criteria exceeds 60% of the potential, and the best-scored, differentiation and message clarity, top out at 52%.

GEO readiness is therefore the weak point of an already weak set.

Does it matter today? ChatGPT referrals to B2B sites multiplied by four in a year according to Demandbase, a 303% rise, and IDC (International Data Corporation) projects that 70% of B2B buyers in the United States will rely on generative AI to discover and evaluate their vendors by 2028. Each will judge the urgency. The measurable point is already here: the channel is growing, and three quarters of the sites we audit are not readable on it.

The thirty-second test

Disable JavaScript in your browser. Reload your homepage.

What stays on screen is, as a first approximation, what an answer engine can read of you. If your message is there but your logos, your figures and your prices have vanished, you know what is left to do, and the fix is a targeted intervention, not a rebuild.

To go beyond the first approximation, you have to measure what we call recoverable message: the share of your positioning that survives a read without JavaScript, extractor included. That is what our audit protocol measures, edge cases included: three extractors, and a check that the promise, the differentiators and the proof come through in the extracted text.

The window is still open. It will not stay open: your competitors are publishing HTML.

Sources

  1. Microsoft Clarity. See Which AI Crawlers Drive Real Website Visits, « AI Scrape-to-Referral Ratio » card, Microsoft Clarity blog, 13 August 2026, reported 19 August. Public example: about 6,000 pages scraped per referred visitor. clarity.microsoft.com
  2. Cloudflare. Crawl-to-referral ratio measured across its network: nearly 70,900 requests per referred visitor for Anthropic’s crawler over June 2025, then an order of magnitude around 4,580 to 1 a year later (Cloudflare Radar). Primary URL to be set at integration.
  3. Vercel and MERJ (web analytics firm). The rise of the AI crawler, 17 December 2024. Server-log study: none of the major AI crawlers render JavaScript, JavaScript (JS) files are fetched without being executed (11.50% of requests for ChatGPT, 23.84% for Claude). vercel.com
  4. FGA Message-Market Fit Observatory, V2 edition. Readability-by-AI-engines sub-criterion at 0.91 out of 5 across 94 full audits; none of the fifteen sub-criteria above 60% of potential; « proof absent from the site » pattern in 61% of 83 detailed audits. Internal Fast Growth Advisors measurements.
  5. Demandbase. Year-on-year tripling of referred visits from ChatGPT to B2B sites (645,000 to 2.6 million visits per month, June 2025 to June 2026). Primary URL to be set at integration.
  6. IDC. Projection: 70% of B2B buyers in the United States will rely on generative AI to discover and evaluate their vendors by 2028. Primary URL to be set at integration.
  7. Fast Growth Advisors. In-house extraction measurement of 29 July 2026, reproducible (cas_limites_realiste.html, test_extracteurs.py), across jusText, Resiliparse and Trafilatura.

FAQ

What does readability by AI engines mean?

It is the share of your positioning that an answer engine can actually read and restate, once it has passed the three stages that separate the URL from the exploited text: the fetch, the server response and the extraction. A site can render perfectly in a browser and still be unreadable to an engine that executes no JavaScript and discards part of the document at extraction. Readability is measured on what the tool keeps, not on what the screen shows.

Why can ChatGPT or Claude ignore content that displays fine in my browser?

Because these engines do not render the page. They request the raw HTML and read what is in it, without executing JavaScript or applying CSS. Content built in the browser, such as a client-rendered React application, arrives empty. A Vercel and MERJ study from December 2024 showed that none of the major AI crawlers render JavaScript. Your browser does that work, which explains the gap between what you see and what an engine keeps.

Is the “disable JavaScript” test enough to know what an AI reads of my site?

It gives a faithful, free first approximation. Disable JavaScript, reload your homepage, and what stays on screen is close to what an engine can read. To go further you have to add the extraction stage, because text present in the HTML can still be discarded by the extractor. The full measurement is what we call recoverable message, the share of positioning that survives a read without JavaScript, extractor included. It says nothing, however, about what an extractor discards once the page has been fetched, which is the second half of the diagnosis.

Do I need to rebuild my site to be readable by AI?

Rarely. In most of the cases we audit, the message already gets through and it is the proof that disappears: logos in a JavaScript carousel, figures animated on scroll, pricing toggles, reviews in a third-party widget. Bringing those elements back into server-rendered HTML text is a targeted intervention, not a reconstruction. A rebuild is only justified when the entire site is client-rendered, the case where extraction recovers nothing at all.