AI search · Reading the evidence
One report says your AI visibility is rising; another barely registers your brand. Before changing your SEO budget, check whether the reports measure the same thing. Different questions, markets and counting rules can produce different scores without either tool being broken.
A score can rise while the answers stay the same
Imagine a service company tracking twenty AI-search questions. Ten ask about the company by name; ten ask for providers without naming it. The company appears in eight branded answers and none of the unbranded answers. Its simple mention rate is eight out of twenty: 40%.
The next report adds ten more branded questions. The company appears in ten of those answers. Nothing changes in the original twenty, but the combined rate becomes eighteen out of thirty: 60%.
| Question set | First report | Next report |
|---|---|---|
| Original 20 questions | 8 mentions / 20 answers | 8 mentions / 20 answers |
| 10 newly added branded questions | Not included | 10 mentions / 10 answers |
| Combined mention rate | 40% | 60% |
The headline improved by twenty percentage points. Discovery in the original unbranded questions did not improve at all. The change came from the sample, not better performance on the questions already being tracked.
This is why a monthly comparison needs a stable core question set. New questions are useful for exploration, but report them separately until you can compare them across equivalent periods. Keep branded and unbranded results distinct: recognition when someone already knows your name is not the same task as being discovered.
“Visibility” is not one standard unit
Current product documentation makes the differences concrete. Ahrefs Brand Radar counts a brand mention when the brand appears at least once in a response. It distinguishes that from a citation of a page or domain. Its impressions use Google search-volume data associated with prompts, and its AI share of voice compares those impressions across tracked brands. Those impressions should not be read as a direct count of people who saw an AI answer. See Ahrefs’ metric definitions.
Semrush describes its 0–100 AI Visibility Score in terms of topic coverage and mention consistency. Its documentation also separates broad discovery reports, brand-performance reporting and custom prompt tracking. Those are different datasets and reporting tasks, not interchangeable views of one universal ranking. See Semrush’s data methodology.
Even inside one product, similarly named metrics can differ. Semrush’s metric reference describes Brand Performance share of voice using mentions, while Prompt Tracking visibility concerns citation positions. A screenshot labelled “visibility” is not enough to identify the measurement.
The practical conclusion is ours: do not average two vendors’ scores, subtract one from the other or treat a change of tool as a continuous performance trend. First identify the exact report, what is counted and what it is counted against. A score of 60 on a 0–100 index does not automatically mean that 60% of potential customers encountered the business.

Ask for a measurement handover, not another screenshot
Before accepting a growth claim, request a short explanation that another analyst could follow. The following is a reporting brief, not a demand for access to a vendor’s proprietary algorithm:
“Name the tool and report, platforms, collection dates, market and language. Show the questions or explain how the sample is built. Separate brand mentions from citations of our website, and explain the denominator or index definition. Identify changes to the question set, brand matching, competitors or collection method since the previous report. Include example answers behind the claimed movement.”
For a company serving Belgium, ask whether the evidence concerns the relevant services and Dutch-, French- or English-language customer questions. A broad English-language global sample may reveal competitors or content ideas. It does not, by itself, establish visibility for a French-speaking buyer looking for a provider near Brussels.
Brand matching deserves a spot check too. A common company name, an old trading name or a product shared with another business can distort an automated count. Read a few counted answers and inspect the cited URLs. Record an incorrect identity or unsupported claim separately from a genuinely useful recommendation.
Repeated collection matters, but repetition does not turn a convenience sample into a census of customer demand. Ahrefs’ discussion of AI tracking limitations explains why one traditional rank-style snapshot is a poor model for variable AI answers. Keep collection conditions consistent and retain the answers, not only the aggregate score.
Let the type of disagreement determine the next action
If the scope differs, reconcile the reports. Align platform, language, market, period and question intent where possible. If the tools cannot produce a comparable subset, keep separate baselines. Do not commission a website rewrite to solve a reporting mismatch.
If comparable results change, inspect the answers. Look for a recurring, commercially relevant pattern: the wrong service attributed to the company, an outdated source, or a relevant customer question your website leaves unanswered. Improve the specific evidence or page involved, then observe the same question set again. A before-and-after change alone does not prove that your edit caused it.
If a score rises but the business outcome is unknown, keep the claim narrow. You have evidence of movement in that measurement—not proof of extra demand, enquiries or revenue. Connect the visibility review to the value of identifiable AI referral traffic without pretending that every exposure produces a trackable click.
The next useful deliverable is a one-page explanation of the report you already receive. If it cannot connect its questions to your services, semantic analysis of customer intent can define a better sample. If the evidence reveals a real visibility gap, SEO and GEO optimisation can turn it into a focused content decision rather than a campaign to inflate a number.
Questions to settle before comparing scores
Is a low AI visibility score proof that my SEO is failing?
No. First inspect the report’s coverage, counting rules and relevance to your customers. A low score in an unsuitable sample cannot establish that your broader SEO performance is failing.
Can I compare scores from two AI visibility tools?
Only when you understand and align the metrics and scope. Otherwise, use each tool as a separate baseline and compare underlying questions and answers rather than treating the scores as the same unit.
Should branded questions be removed from the report?
No. They help assess how AI describes a business that a customer already knows. Report them separately from unbranded discovery questions so a change in the mix does not hide what improved.