Skip to content
Our own measurement, 2026-09-29

llms.txt: who makes it and who reads it

The llms.txt file is held to be a way of explaining yourself to AI assistants. We measured both sides: how many sites have made one, and how many times anybody fetched it.

Short answer: the file exists on 17.2% of the sites checked, it is made carefully and by hand, and the visitors coming for it are SEO scanners, technology directories and people in browsers. There was not a single AI agent among them.

How many sites have made one

The sample is 285 development studio domains from a web development ranking. The check runs on the same code that powers our free tool.

Sites checked 285
A genuine llms.txt 49 (17.2%)
Serving an HTML page at that address 13

The trap we fell into ourselves. The first measurement counted 62 sites with a file, and that figure (21.8%) made it onto our tool page. A check for "status 200 and a non-empty response" counts an ordinary HTML 404 page that the server returns with status 200. After filtering, 49 remained. If you measure this yourself, discard responses starting with <!doctype.

The files are made seriously

This is the most surprising part of the measurement. It would have been easy to assume the file gets dropped in just in case and filled with rubbish. Nothing of the sort.

What we checked Of 49 files
Fully follows the specification 37 (75.5%)
Has a first-level heading 45 (91.8%)
Has a short description as a blockquote 37 (75.5%)
Split into sections 45 (91.8%)
No links inside at all 10 (20.4%)

The median is 8,519 bytes, 85 lines and 19 links. That is not a formality but half an hour to an hour of work per file.

And the decisive detail: not a single identical file was found on two different domains. We compared content by fingerprint with whitespace normalised. So this is not the output of a shared plugin — every file was written by hand.

Who comes for it

The second half of the measurement is our own access logs over 15 days, 23,957 lines. Total requests to /llms.txt: 27, of which 8 were our own, from the tool and from checks. That leaves 19.

Who fetched it Requests
browser (a person) 8
yacybot 3
BuiltWith 2
agentcensus-crawler 1
SERankingBacklinksBot 1
Baiduspider 1
DropMonBot 1
Dataprovider.com 1
SeoScannerBot 1
Of those, AI agents 0

We looked for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot, Applebot-Extended, Bytespider and several more. Not one of them came for the file, not once.

For comparison, in the same window and the same log the sitemap was requested 656 times. The difference is not a multiple but two orders of magnitude.

What these numbers do not prove

  • The logs of one small site. 15 days and 23,957 lines. On a site with a different audience the picture may differ. This is an observation, not a law.
  • We cannot see robots.txt. Its location in nginx had access_log off, so requests to it never reached the log at all. We found that while preparing this material. So we compare against the sitemap rather than robots.txt, though comparing with robots.txt would be fairer.
  • No request does not mean no use. An assistant may know the contents of the file from a search index without fetching it directly. We measure requests, not influence.
  • The sample is biased towards the strong end. These are agencies from a ranking, meaning people who work on their own presence. Market-wide the share of files will be lower.
  • Tomorrow this may change. The specification is less than two years old. If agents start reading the file, the measurement goes stale — and we will repeat it.

What to do about it

If you do not have the file, do not rush. Half an hour is better spent on what agents actually fetch: the sitemap, access in robots.txt, and text visible without JavaScript. By our market measurement, 12.9% of sites have no Sitemap directive, and 56.3% have no fragment a model could quote whole. That is cheaper and it definitely works.

If the file already exists, do not delete it. It does no harm, costs nothing to maintain, and if agents do come for it you will be ready. Just do not expect traffic from it today.

And do not trust reports presenting a rise in llms.txt requests as interest from assistants. Judging by our log, those are technology directories and SEO scanners — they crawl the file precisely because so much is written about it.

Common questions

What is llms.txt?

llms.txt is a text file in the root of a site which, in the intent of the llmstxt.org specification, describes the structure of the site to a language model: a heading, a short description and sections with links to the main pages. The idea is the same as robots.txt and sitemap.xml, except the addressee is an assistant rather than a search crawler.

Does a site need llms.txt?

By our measurements, today almost not. Over 15 days the file was fetched from our site 27 times, and not once was it an AI agent: the visitors were SEO scanners, technology directories and people in browsers. The sitemap was requested 656 times in the same period. The file does no harm and costs half an hour, but there is no basis for treating it as a route into assistant answers.

How many sites have made one?

In our sample of 285 development studio websites a genuine file was found on 49, which is 17.2%. An important correction: a naive check for "status 200 and a non-empty response" gives 62 sites, because 13 of them serve their own HTML 404 page at that address. We made that mistake ourselves once.

Are the files made seriously or for show?

Seriously. 75.5% of the files follow the llmstxt.org specification: a heading, sections and links in markdown. The median is 8,519 bytes and 19 links. Not a single identical file was found on two different domains, which means this is handwork rather than the output of a plugin.

How do I check this on my own site?

In the web server access logs: look for lines with GET /llms.txt and check the user agent. Two traps. First, the shared access log may contain only the http-to-https redirect host — you need the log of the specific virtual host. Second, vulnerability scanners impersonate AI bot user agents, so count successful responses only.

How to reproduce this

Both halves of the measurement run as scripts rather than by hand, and the sample is stored as a file — otherwise the result cannot be repeated, and there would be no reason to believe it. The site half: assemble the sample, then run the content check across it. The log half: take the access logs and break requests down by user agent, counting successful responses only.

You can check your own site by the same rules in the free tool — it asks for nothing and shows your result against 279 studio websites.