The AI visibility measurement gap
Something happens dozens of times a week that your analytics will never record. Someone asks an assistant to recommend a supplier in your category. Your name is either in the answer or it is not. If it is, you may get a branded search a week later, or an enquiry that mentions being recommended, or nothing traceable at all.
This is now a real commercial channel with almost no native instrumentation. The click-through rate from AI answers is a fraction of traditional search — AI Mode refers at roughly 1.6–2.5% of queries against 17–19% for conventional Google results — so judging your AI presence by referral traffic is like judging billboard performance by counting people who phone the number.
The good news is that it is testable directly, and you do not need to buy anything to start.
Build the prompt set first
The whole method depends on asking the right questions, so this step deserves an hour of real thought rather than five minutes.
Write twenty to thirty prompts in the language a genuine buyer would use, split across four types:
- Category recommendation. "Best web design agencies for D2C brands in India." "Who should I hire to build a Shopify store?"
- Problem-first. "My website gets traffic but no leads, who can help?" These matter more than category prompts because they are how people actually ask.
- Comparison. "X versus Y for a B2B SaaS company." Include competitors by name.
- Direct brand. "What is [your brand] and what do they do?" This tests accuracy rather than presence, and the results are frequently alarming.
Keep the list fixed. The value comes from running the same prompts repeatedly over months, and a prompt set you keep improving produces data you cannot compare against itself.
Test across assistants, not one
This is the step people skip, and it invalidates most casual conclusions.
Different assistants draw on genuinely different source sets. An analysis of around 680 million citations found only about 11% of domains were cited by both ChatGPT and Perplexity. Separate research measured brand citation rates differing by as much as 46 times between platforms. Being absent from one assistant tells you close to nothing about the others.
Run the set through at least ChatGPT, Perplexity, Google's AI results, and one more that matters in your market. Use a logged-out or fresh session where you can, because personalisation and memory will otherwise show you a flattering result built from your own history.
Record four things, every time
A spreadsheet with one row per prompt per assistant per month. For each, capture:
- Were you mentioned at all? Yes or no. This is the headline metric.
- Where in the answer? Named first, in a list of three, or buried at position nine. Position matters roughly as much as it does in search results.
- Was it accurate? Wrong services, wrong location, wrong pricing, a competitor's work attributed to you. This is the finding that most often needs immediate action.
- What sources were cited? The most useful column by far, because it tells you where the answer came from and therefore where the work is.
That fourth column is the whole strategy in disguise. If the same three directories and two industry publications appear across every answer in your category, you have just been handed your outreach list.
What the sources column usually reveals
The consistent finding when businesses run this exercise is that assistants lean far more on third-party sources than on the brand's own website. Industry publications, directories, review platforms, forum and community discussion, and comparison articles on sites you do not control show up repeatedly. Your own site tends to be cited for factual specifics — what you do, where you are, what you charge — rather than for evaluative claims about how good you are.
Which makes the division of labour clear. Your site's job is to be accurate, specific, and crawlable so that when it is consulted the facts are right. Your citation profile is built somewhere else, through the unglamorous work of being present and well-regarded in the places your category gets discussed. The tactics for that are in our guide to generative engine optimisation.
Does llms.txt help?
It comes up in every conversation about this, so: there is no good evidence it works. Research through mid-2025 found the major LLM crawlers were not fetching llms.txt, and no major assistant has committed to honouring it.
It costs almost nothing to publish, so treat it as a cheap option on a future that may or may not arrive. What genuinely matters and is frequently broken is the opposite: check that your robots.txt, WAF, and bot-protection rules are not blocking the assistant crawlers you want citing you. We have seen sites invisible to AI search for no reason other than an over-aggressive firewall rule nobody remembered adding — the same class of problem as the ones in our indexing checklist.
When to buy a tool
Manual testing scales to about thirty prompts across four assistants once a month. Past that it becomes a job nobody does consistently, which is worse than not doing it.
Buy a tool when you need daily rather than monthly tracking, coverage across several markets or languages, competitor benchmarking, or a report someone else will read. Before that point, the spreadsheet is not a compromise — it is better, because you read every answer yourself and notice the things a dashboard reduces to a score.
Whichever route you take, pair it with branded search volume in Search Console. If AI visibility is working, branded search rises even while unbranded traffic falls, which is the pattern we unpack in the zero-click reality. If you would like us to run the first pass for your category, tell us who you compete with.