Can ChatGPT see your site
Assistants do not run scripts and do not browse a site — they take what the server returned immediately. This check shows whether you let their agents in, whether the text is visible without JavaScript, and whether it contains a fragment that can be quoted whole.
The tool reads three public addresses: the home page, /robots.txt
and /llms.txt. It changes nothing, stores nothing and asks
for no access. It was built from an analysis of our own server logs —
what came out of that is below.
What the market looks like
For "three out of six" to mean anything you need a baseline. We ran the same check over 279 development studio websites — a sample from a web development ranking, 285 domains, of which 279 responded. Measured in 2026-09-29.
| Check | Pass |
|---|---|
| Access for AI agents | 99% (276 of 279) |
| Text visible without JavaScript | 94% (263 of 279) |
| A fragment worth quoting | 44% (122 of 279) |
| Title and description | 94% (261 of 279) |
| Schema.org markup | 44% (124 of 279) |
| Sitemap in robots.txt | 87% (243 of 279) |
The market did SEO but not quotability
Title and description are present on 94%, an H1 on
84%. But a fragment a model can
take into an answer whole exists on only 44%.
FAQPage markup, which is what
answers to questions get assembled from, is on
4%.
People make llms.txt and nobody reads it
llms.txt exists on
18% of the
sites checked, and it is made seriously — three quarters of the files
follow the specification and were written by hand. Yet across fifteen
days of our own logs not one AI agent requested it: the visitors were SEO scanners and people in browsers.
Almost nobody blocks access
3 sites out of 279 block at least one AI agent, and 1 block all of them. If someone tells you that "opening your site to AI" is the first step — it is probably already open. The barrier is not robots.txt.
The text is usually not the problem
The median amount of text without JavaScript is 6,815 characters, so the content is visible to an assistant almost everywhere. But the median longest continuous block is 694 characters: there is text, and often no complete thought to quote.
What is wrong with these numbers — read this before citing them
- The sample is biased towards the strong end. These are agencies from a ranking, meaning people who work on their own presence. Market-wide the numbers will be worse, not better.
- We invented the checks. Our own site passes all six, which is unsurprising, since we wrote them. That is not proof we are right; it is a reason to read the justification for each one.
- A repeat run differs by about a point. Some sites do not respond first time. A difference of one or two points cannot be treated as meaningful.
- Only the home page is checked. We look at whether an assistant can see the company at all, not every section.
- The measurement is reproducible. The sample is stored as a file and the measurement runs on the same code that powers this page. Our earlier measurement was a one-off, and its numbers had to be thrown away when an error turned up in the metric.
Where these checks came from
We went through fifteen days of our own access logs and got a picture we had not found in any article on the subject.
- AI bots are busier than Googlebot. 431 successful fetches against 84. Most often ClaudeBot and PerplexityBot.
- Real people arrive from ChatGPT. Not many, but they read: three minutes twenty-three on site, zero bounce.
- Nobody reads llms.txt. Over the same fifteen days no agent requested it. The sitemap was fetched 168 times.
- The numbers are easy to inflate sevenfold. Vulnerability scanners impersonate AI bot agents: of 3,015 such lines,
435 were genuine and the rest were requests to
/fetchand/.envreturning 404.
Hence the order of the checks: access and text without scripts first, then quotability, and llms.txt last with the caveat that it decides nothing yet.
The full analysis, with the numbers, the commands and both traps, is in the article on measuring ChatGPT traffic from server logs. It also explains how to repeat the measurement yourself.
Common questions
What does visibility for AI search actually mean?
Whether your page can reach an assistant's answer — ChatGPT, Perplexity and the rest. Three things decide it: whether robots.txt lets their agents in, whether the text is visible without JavaScript, and whether there is a fragment that can be quoted whole. Assistants do not run scripts and do not browse a site; they take what the server returned.
Does a site need an llms.txt file?
By our measurements, almost not. We went through fifteen days of our own access logs: not a single agent requested llms.txt, while ordinary pages were fetched in the hundreds and the sitemap 168 times. The file does no harm, but there is no basis today for treating it as a working channel.
How is AI traffic different from search traffic?
In volume and in behaviour. Over fifteen days AI bots made 431 successful fetches on our site against 84 by Googlebot, and the humans arriving from ChatGPT spent three minutes twenty-three on the site with a zero bounce rate. There are few of them, but they read rather than bounce.
How do I count AI traffic on my own site?
In the web server access logs: look for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot in the user agent, and separately for human referrals tagged utm_source=chatgpt.com. One trap matters: vulnerability scanners impersonate AI bot agents. Of 3,015 such lines in our logs only 435 were genuine; the rest were requests to /fetch and /.env returning 404. Count successful responses only.
What is my result compared against?
Against a measurement across 279 development studio websites: we ran this same check over a sample of agencies from a web development ranking in 2026-09-29. Access for AI agents turned out to be open almost everywhere (3 sites out of 279 block at least one), title and description are present on 94%, and a fragment a model can quote whole on only 44%. A separate surprise: 18% of sites have made an llms.txt, although no agent in our logs ever requested one. The sample is biased towards the strong end — these are agencies from a ranking — so market-wide the numbers will be worse.
Does the check change anything on my site?
No. The tool reads three public addresses: the home page, /robots.txt and /llms.txt — exactly what any visitor sees. It sends nothing, stores nothing and needs no access.
We are MZR Digital, a custom development studio. We build AI assistants on local LLMs and RAG, B2B platforms and replacements for foreign ERP systems. This tool is a side effect of working on our own site, and it is free with no conditions.