AI brand perception is measured by putting questions to AI systems and scoring what comes back. There are two broad approaches. Visibility monitoring tracks which brands appear in AI answers to a chosen set of queries, how often, and which sources are cited. Controlled measurement puts the same standardised test to every brand and every model, so the readings can be compared brand with brand, model with model and week with week. They answer different questions, and both can be useful.
Two questions, two methods
The simplest way to tell the approaches apart is by the question each one asks.
Observational measurement asks: what is happening in AI conversations? Controlled measurement asks: what happens when AI systems are all given the same test?
Neither question is the wrong one. They measure different things, and the right choice turns on what you need to know.
How visibility monitoring works
Visibility monitoring starts from a set of queries and records what AI systems say in response. By approach, the queries may come from real users, from panels built by analysts, or from prompts the customer chooses. The answers are then counted: how often a brand is mentioned, how often it is named first, its share of voice against competitors, and which web sources the answer cites. Some approaches add a sentiment reading.
This is useful. It shows where a brand turns up, which sources AI leans on, and how that changes as content and coverage change. For a team working on search and content, it is a direct view of the answers people may be seeing.
Because it starts from a chosen set of queries, the results are shaped by that set. Which questions are asked, in which language, from which market, shapes which brands and attributes are measured. Where queries come from real-world behaviour, the characteristics of the people asking become part of the measurement. That is not a flaw; it is what observational measurement is for.
How controlled measurement works
Controlled measurement takes the opposite starting point. Instead of following the questions people happen to ask, it defines one test and applies it to everything. The Pitot Index® works this way. That Index score is an AI Brand Rating: a measure of AI Brand Understanding (AIBU), how AI models understand a brand in the first place. In practice that means:
- One standardised test. Every brand in an industry goes through the same questions, the same evaluation criteria and the same scoring. The principle is simple: change the brand or the AI model, not the rules of the test.
- Unprompted. The models are never told which brands to discuss. They bring up, rank and characterise brands on their own, and the Index records what they bring up.
- Several models, pinned. No single model represents AI as a whole. The Index runs on a set of pinned AI models, currently nine, from Western, Chinese and Arabic-first developers. The model actually served is recorded separately from the model requested. We check model versions every cycle and log any change we find, with the date we found it.
- A rating, not a count. Each brand is assessed on five dimensions (Authority, Reputation, Expertise Visibility, Signal Consistency and Context), and its score combines them. Signal Consistency is derived from how far the models agree.
- Automated, and weekly. Scores are produced automatically, and no person adjusts a score to suit a brand. Every measured brand is re-measured every week, on the same instrument, and the record is published as each cycle closes.
- Versioned and documented. The same version of the test is used from cycle to cycle, and any change is versioned and dated. Because the test is standardised, it can measure its own week-to-week noise, so a reported move can be read against the instrument's own variation.
The aim is comparability: brand against brand, model against model, and one week against the next.
Controlled measurement does not use private conversations. Pitot Group does not access, purchase or analyse people's chats with AI systems, and the Index needs no personal data. The prompts are written for the research and are independent of any individual user. The questions themselves, and how they are scored, are not published: released in full, they would let anyone run a lookalike whose numbers look comparable but are not.
The Pitot Index never names the brands under test. Pulse and the Audit do. What the three share is where the questions come from: none is lifted from what people actually ask AI, and every one is written for the research. The starting point sits as far from any single person as it can, to keep the result as neutral as we can make it.
Studying real queries measures something else, and measures it well: what a defined group of people asks, how often, and what the models they happen to use say back. The pool is set by who is asking, and each real query carries the framing of the person who asked it. Every model also brings its own perspective, shaped by its training data, which is why the Pitot Index reads several side by side. The questions belong to the research, and to no one else.
What controlled measurement does not tell you
It is important to be plain about the limits.
A controlled score is an automated reading of what AI models output under a standardised test. It is not a measure of what consumers do, how often real users are shown a brand, or how often a brand is actually recommended in everyday use. It is not a statement about a brand's quality or conduct. Search rankings, AI mentions, cited sources and sales are separate measures and should be tracked separately.
Controlled measurement also has two modes, and they are never mixed. The Pitot Index is unprompted. Commissioned assessments of a single brand, such as Pitot Group's Audit, are prompted: the brand is named and the instrument examines how the models characterise it when asked directly. The same brand can score very differently under the two, so an unprompted ranking is never compared with a prompted assessment.
Which do you need?
It turns on the question you are trying to answer.
If you already run GEO or AEO, you are used to questions about mentions, share of voice and sentiment. Those read what appears in AI answers. AIBU asks a prior question: how do the models understand the brand? The Pitot Index score is Pitot Group's AI Brand Rating of that understanding.
- "Where does my brand show up in AI answers, and which sources are cited?" That is a visibility question. Monitoring a relevant set of queries answers it directly.
- "How does AI rate my brand against its peers, and is that changing?" That is a comparability question. It needs a test that stays the same for every brand and every week.
- "Do the models agree about us?" That needs several models read side by side, showing where they agree and where they diverge.
GEO and AEO remain useful for what they are built to do. The risk is sequencing. In our view, constant optimisation of mentions, share of voice and sentiment without an AIBU reading can add inconsistent content that tries to force-correct those readings, and over time the brand image the models hold can drift. The Pitot Index does not replace GEO or AEO; it gives a diagnosis those programmes can build on.
Many teams will want both kinds of reading. What matters is not to treat one as the other: a brand can be mentioned often and still be rated below its competitors, because being visible to AI and being recommended by AI are not the same thing.
For how Pitot Group's work relates to SEO, GEO and PR, see our questions page.
Whatever the method, no measurement can guarantee a place in an AI answer. What it can do is show, consistently, where a brand stands.
Sources and method
- AI Brand Understanding and AI Brand Rating are Pitot Group's terms.
- Observational versus controlled, "change the brand or the AI model — not the rules of the test", unprompted and prompted, no private conversations, what is not published, noise measured directly, visible versus recommended: pitotindex.com/methodology (read 8 Oct 2026).
- "An automated reading … not our opinion of the brand, and not a statement about its actual quality or conduct": pitotindex.com/faq.
- "Scores are produced automatically, and no person adjusts a score to suit a brand": pitotgroup.com/measure/how-we-measure/.
- "Currently nine … Western, Chinese and Arabic-first developers", "the model actually served is recorded separately", and the version-check line: Legal-cleared texts C1–C3 and Q2 (3 Oct 2026).
- The description of visibility monitoring is generic. It follows the no-names public version (Version B) of
launch/explainer-controlled-vs-observational.md, which Legal reviewed on 29 Sep 2026. No provider is named or described individually. - The brief's noise figures ("fewer than one case in thirty…") are not used: the live methodology page no longer states them.
- No prompts, label sets, score mappings or raw data are disclosed.