Buyer-focused prompts, platform variation and descriptive metrics distinguish mentions, citations and links from meaningful business outcomes, according to this writer
AI is changing some search behaviours and creating new questions about how brands measure their visibility in answer-driven interfaces.
When was the last time you Googled something and actually clicked on a webpage? For some queries, users may now receive an answer from Google AI Overviews or turn directly to a chatbot. This may be a relatively small change from the consumer side, but it is changing some of the ways brands are discovered online.
As a result, digital strategies are evolving. Where SEO has traditionally focused heavily on ranking in search results, some teams are also seeking to understand whether their brands are cited or mentioned inside AI-generated answers.
The old and emerging landscapes can operate with different goals:
- AEO (Answer Engine Optimisation) generally concerns making content easier to surface in direct-answer formats, including featured snippets and structured search features.
- GEO (Generative Engine Optimisation) is commonly used to describe efforts to create authoritative, in-depth content that generative AI systems may summarise, cite or attribute within a conversational response.
A search strategy may include both conventional SEO practices and efforts to understand a brand’s visibility in answer-oriented and generative interfaces.
Why prompts matter
Traditional search results can vary by location, time, device and user context. AI-generated answers can introduce additional variation, as outputs may differ by model version, conversation context and prompt wording. There is therefore no single fixed “rank” that translates neatly across AI interfaces, which makes prompt selection an important consideration in GEO and AEO dashboards.
In our client benchmarking, we have seen that generic category prompts can produce results that look favourable but may not reflect the questions real buyers ask. A brand may appear strongly for broad prompts while being less visible for queries linked to specific buying needs.
Running more prompts does not necessarily resolve that issue. Greater volume can make flawed measurement more expensive if the prompt set does not reflect the audience or decision being examined.
One approach is to build prompts around buyer language and decision stages. In my approach, a four-part methodology is used:
- Segment sets the market context
- Persona identifies who is asking
- Intent defines what they are trying to achieve
- Variable isolates a single factor to test, such as “fastest” versus “cheapest”
A manageable set of prompts built this way may be more useful than a much larger set of broad generic queries, although the appropriate number will vary according to the market, audience and product category. Once the queries have been defined, teams can consider two broad measurement layers.
Measuring AI visibility
The first is primary metrics coming directly from tracking platforms. They can provide a baseline view of how a brand appears across a defined set of prompts.
- Share of voice is one measure. It refers to the proportion of AI-generated answers in a defined category and prompt set that reference one brand relative to competitors.
- Visibility percentage measures the proportion of targeted queries in which a brand is mentioned or cited by the AI platforms being monitored.
- Mention frequency refers to how often a particular entity, brand or term appears in a defined set of AI responses.
These measures are descriptive. Their usefulness depends on the selected platforms, prompt set, competitor group, time period and the way a provider identifies mentions, citations and brand entities.
Secondary metrics can be used to examine consistency and variation behind those surface numbers. For example, a run-length measure can record how many consecutive observation periods a brand is visible for a topic. A short run may indicate that results are variable within the test, while a longer run may indicate greater consistency. It does not, on its own, establish a durable association between a brand and a topic.
- Shannon entropy can be used to describe how evenly visibility is distributed across brands within a defined topic and sample. Low entropy indicates that mentions are concentrated among fewer brands, while high entropy indicates a more evenly distributed set of mentions. Neither result alone establishes whether a category is commercially attractive or whether a brand’s appearance is valuable.
- Kullback–Leibler divergence can be used to compare the distribution of results on one AI platform with a chosen reference distribution across tracked platforms. It may help identify platform-specific variation, but strong visibility on one product is not necessarily invalid or less valuable than visibility across several products.
Modern search measurement is therefore different from traditional rank tracking in some respects, but AI visibility should not be treated as a replacement for established search, traffic and conversion measures.
By using clearly defined prompts and distinguishing between mentions, citations, links and business outcomes, brands can develop a more useful view of how they appear in AI-generated answers. Any such measurement should be treated as directional and assessed alongside referral traffic, enquiries, conversions and other outcomes relevant to the organisation.




