Free test · AI assistants
Is your site open to AI?
ChatGPT, Claude, Perplexity, Copilot and Google only cite what they can read. An overly broad robots.txt, a firewall set once then forgotten, a page that exists only in JavaScript: the door can close without anyone knowing. Paste your address, and we will look.
We read your robots.txt, your home page as a browser would, then the same page on behalf of eight AI bots. Eleven reads in all, forgotten as soon as they are done. We read public addresses, and those alone.
What the test looks at, and what it proves
Three reads, three different strengths of proof. The result always says which one spoke.
1. Your robots.txt — certain
The file is public and reading it is standardised (RFC 9309). We read it as the bots read it: the group naming the bot prevails over the general group, the longest rule wins, and in a tie the allow wins. A robots.txt that answers with a server error closes the whole site to a well-behaved bot: we flag it.
2. The response given to the bot — likely
We request your home page as a browser, then on behalf of GPTBot, ClaudeBot, PerplexityBot and the others. If the browser receives the page and the bot a refusal, a firewall is filtering AI bots. We say “likely” and not “certain”: the real bot comes from its publisher’s addresses, and some firewalls treat it differently. Googlebot and Bingbot are judged on robots.txt alone: firewalls block impostors on principle, and you deserve an accurate verdict.
3. What the page contains as it is served
Assistant bots read the page as it arrives. They do not run JavaScript. A page assembled in the browser reaches them almost empty: we count the words actually served. We also note
noindexandnosnippet, which remove the page from the engines the assistants rely on.
The bots, one by one
Each publisher sends out several bots, and each one has its own role. The bots of search decide whether you can be found, and therefore cited. The bots of on-demand reading open your page when a user hands it to the assistant. The bots oftraining collect data to train the models: closing them is a legitimate choice, and keeps you in every answer.
| Bot | For | Role | If you close it |
|---|---|---|---|
| OAI-SearchBot | ChatGPT | Search | Your pages no longer appear in ChatGPT's answers with search. |
| ChatGPT-User | ChatGPT | On-demand fetching | ChatGPT can no longer open your page when a user asks it to. |
| GPTBot | OpenAI | Training | Your pages are no longer used to train the models. Search remains unaffected. |
| Claude-SearchBot | Claude | Search | Your pages no longer appear in Claude's answers with search. |
| Claude-User | Claude | On-demand fetching | Claude can no longer open your page at a user's request. |
| ClaudeBot | Anthropic | Training | Your pages are no longer used to train the models. Search remains unaffected. |
| PerplexityBot | Perplexity | Search | Your pages are removed from Perplexity's index. |
| Perplexity-User | Perplexity | On-demand fetching | Perplexity can no longer open your page on demand. |
| Bingbot | Copilot | Search | You're opting out of Bing — and everything that relies on its index, including Copilot. |
| Googlebot | Search | You're opting out of Google, including AI Overviews. | |
| Google-Extended | Training | Gemini no longer uses your pages for training. Google Search remains intact. | |
| MistralAI-User | Le Chat | On-demand fetching | Mistral's Le Chat can no longer open your page on demand. |
| CCBot | Common Crawl | Training | Your pages are leaving the open archive that many models rely on. |
Close off training, stay in the answers
This is the most common confusion: you want your texts not to train the models, so you block "the AIs" across the board, and you drop out of their answers too. The two can be separated. This robots.txt closes off training at OpenAI, Anthropic, Google and Common Crawl, and leaves search open everywhere:
# Fermer l'entraînement, garder la recherche ouverte
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
# Tous les autres robots, dont ceux de la recherche
User-agent: *
Allow: /Multiple lines User-agent in a row form a single group; a bot that has its own group ignores the group *. Run your site through the test after every change: one line too many is enough to close off search.
The firewall: the door that closes silently
A robots.txt gets read; a firewall does not. Cloudflare, certain security plugins and certain hosts offer to block "AI bots" with a single setting. That setting returns 403 to the bots and leaves robots.txt untouched: the site looks open to whoever reads the file, and closed to whoever knocks at the door. That is exactly what the test's second reading brings to light.
At Cloudflare, the setting is in the domain dashboard, under the section devoted to AI bots. It lets you handle each bot separately: close off training, keep search. Cloudflare can also manage your robots.txt for you and write content signals into it (Content-Signal: search=yes, ai-train=no); the test reads them and shows them to you.
Open, then cited: two steps
An open door is the condition, not the result. Assistants cite the pages they find in their search indexes, and there they first find the pages that already rank well. This test tells you whether the way is clear; who actually comes through to your site is measured elsewhere.
This is measured at the network edge, where the requests arrive: which bot read which page, and when, each read certified by the address ranges that publishers make public. That is what GEO SEO Records. The guide Getting cited by AI assistants details what makes a page citable once the door is open.