Does ChatGPT send anyone to your site? Measure it in the access logs
Neither analytics nor Search Console answers whether AI bots visit you. Only your access logs do. Real numbers: 431 fetches by AI bots against 84 by Googlebot, and two traps that inflate the result sevenfold.
Short answer: you can only check this in your web server’s access logs. Analytics platforms do not count bots at all, Google Search Console shows Google, Yandex Webmaster shows Yandex. You need to look for the agents GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, and for human referrals the marker utm_source=chatgpt.com. And count only successful responses, or the number is inflated sevenfold.
We went through fifteen days of our own logs. Below is both the method and the numbers, including the uncomfortable ones.
Why no dashboard answers this question
Tool by tool, because people look in the wrong place:
- JavaScript analytics counts people, not robots. The counter is a script on the page, and AI bots do not execute scripts. They are not there and never will be.
- Google Search Console shows Googlebot only, and only Google’s results.
- Yandex Webmaster shows Yandex only.
- Bing Webmaster Tools has an AI Performance section now, but it counts citations in Microsoft products — Copilot and partners. That has nothing to do with ChatGPT.
Which leaves the server. It records every request together with the User-Agent and Referer, and it is the only place where both the bots and the human referrals are visible.
Trap one: there are usually two logs, and the obvious one is the wrong one
The first thing we tripped over ourselves. A typical nginx setup has a shared /var/log/nginx/access.log plus a separate log for the virtual host.
Ours turned out to be the log of the redirect host (http → https): 50,630 entries, 100% of them status 301 and not a single 200. Drawing conclusions from it is meaningless — those are not visits, they are redirects and scanner noise.
The real site log was named after the site. Before counting anything, check that the file contains any 200s at all:
awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head
If the top of that list is all 301, you are reading the wrong file.
Trap two: vulnerability scanners impersonate AI bot agents
This is the main reason not to trust other people’s numbers on this topic.
Our log contained 3,015 lines with agents like GPTBot and ClaudeBot. A nice figure. But looking at the response codes, 2,572 of them were 404s, and here is what they were requesting:
/fetch /proxy /.env /@fs/etc/passwd /api/webhook /actuator
Those are vulnerability scanners wearing a popular User-Agent to slip past simple filters. Genuine requests: 435 — seven times fewer.
We saw the same trick with “Googlebot”: 35,703 lines from 21 addresses, all status 301, all hitting /.env and /.git/index.
The rule: count successful responses only (200 and 304). It also helps to look at the addresses — real OpenAI and Perplexity bots come from recognisable ranges, not from a single host hammering /.env.
What fifteen days of our logs showed
The period is 5–19 September 2026. Successful responses only.
Bots that crawl on their own:
| Agent | Fetches |
|---|---|
| ClaudeBot | 153 |
| PerplexityBot | 104 |
| GPTBot | 39 |
| OAI-SearchBot | 31 |
| Amazonbot | 15 |
| meta-externalagent | 1 |
| Total | 343 |
Bots that opened a page because a human was asking an assistant right then: ChatGPT-User — 81, Claude-User — 6, Perplexity-User — 1. That is 88 requests, around six a day.
The difference between the two groups matters. The first is crawling for an index. The second is already a conversation with a live person: somebody asked, and the assistant went to read your page in order to answer.
For comparison, conventional search engines over the same fifteen days: Bingbot — 218, Yandex — 201, Googlebot — 84.
So AI bots fetched our pages five times more often than Googlebot did — 431 against 84.
Actual humans from ChatGPT: three visits worth discussing
ChatGPT tags referrals with utm_source=chatgpt.com, so they show up both in the logs and in analytics.
Over fifteen days there were three — on 7, 9 and 14 September. They landed on the piece about what a RAG system costs, on the homepage, and on the terms page.
Analytics showed something the logs do not: 3 minutes 23 seconds on site and a zero bounce rate. For comparison, five visitors from Google over the same period produced five page views between them — a look and a leave.
And one more number we did not enjoy: by unique visitors over fifteen days, Google 5, ChatGPT 4, DuckDuckGo 2, Yandex 1.
What these numbers do NOT mean
We have to be honest here, otherwise this turns into an advert instead of an analysis.
Three visits is not much. The site does not have a lot of traffic yet, and you cannot say “ChatGPT brings clients” on numbers like these. The value is not in the volume, it is in two things: the channel exists and it is measurable, and the ratios — 431 against 84, three minutes against twenty seconds — do not depend on volume.
“Bing’s index feeds ChatGPT” did not hold up for us. It is a popular claim and we repeated it ourselves. After connecting Bing Webmaster Tools it turned out that the AI Performance section showed zero citations through Microsoft Copilots over three months. Meanwhile ChatGPT’s own crawlers visit us dozens of times. The conclusion: these are separate channels, and being in Bing tells you nothing about your visibility to ChatGPT.
Not a single agent requested llms.txt in fifteen days. Not once. Over the same period AI crawlers fetched sitemap-index.xml 168 times — that is their main entry point. The file does no harm, but there is no basis today for treating it as a working channel: plain HTML and the sitemap are what decide it.
How to repeat this on your own site
The minimal version is one command. Replace the path with your log:
grep -E 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|PerplexityBot' \
/var/log/nginx/site.access.log \
| awk '$9 == 200 || $9 == 304' \
| grep -oE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|PerplexityBot' \
| sort | uniq -c | sort -rn
The $9 == 200 filter is the protection against that sevenfold error.
Human referrals are found separately:
grep -c 'utm_source=chatgpt' /var/log/nginx/site.access.log
If your logs are rotated and some are gzipped, a small script that groups agents, discards the fakes and shows which pages were fetched is more convenient. We wrote one for ourselves; it reads plain files and .gz alike.
It is also worth checking whether you let these bots in at all. If robots.txt has Disallow: / for GPTBot, there will be no channel regardless of how good your text is.
What to do about it
Three practical conclusions from our own data.
Aim at the homepage. Of the 88 “live question” requests, 75 hit the homepage. People do not ask an assistant about your article, they ask about your company: who are they, what do they do, what does it cost, where are they. So short self-contained answers belong there first, not deep in the blog.
Write about money and timelines. Of the internal pages, assistants most often opened two: what a RAG system costs and the budget for a local LLM. People ask assistants “how much does this cost”, not “what is RAG”.
Measure regularly, not once. The channel is new and it moves. We now take these numbers every two weeks alongside the query exports from the search dashboards.
We build AI assistants on local LLMs and RAG over corporate documents. This analysis was a side effect of working on our own site: first we needed to answer whether the channel works at all, then it turned out no dashboard has the answer.
Facing a similar problem?
Tell us what you are building. We will walk through the architecture and give you a budget range in one call.