Article

How to measure whether an AI assistant mentions your brand

The prompt matrix, the pinned snapshots, and the arithmetic — plus the four ways this measurement lies to you.

Flexsent Labs · 18 August 2026 · 3 min read · 3 sources

Measurement here is simple to describe and easy to get wrong. You need a fixed set of prompts, pinned model snapshots, and a rule for what counts as a mention — all three decided before you look at a single answer.

The matrix

Write prompts a real buyer would type. Not keywords: questions. “What should I use to monitor brand mentions in ChatGPT?” is a prompt. “AI visibility tool” is a search query, and nobody talks to an assistant that way.

Version the set and freeze it. Never edit a prompt in place — a changed prompt gets a new id and a new matrix version, or your time series quietly stops meaning anything while continuing to produce numbers.

Cover the shapes buyers actually use: direct recommendation, comparison against a named competitor, category definition, and problem-first framing where your category is never mentioned. The last one is usually where brands discover they are invisible.

The pinning

Run against pinned model snapshots, not “latest”. A provider silently upgrading a model underneath you produces a step change indistinguishable from a change you caused. Record the snapshot identifier with every run. If your tooling cannot pin, your measurements are not comparable across months and should not be plotted as a line.

The arithmetic

Count the answers naming you, divide by the answers you ran. Report both integers — never the percentage alone. 3/5 and 300/500 are the same percentage and not the same finding, and the difference matters most exactly when the number is moving.

The four ways this lies to you

Non-determinism. The same prompt against the same snapshot does not return the same answer. Run each prompt several times and report the rate, not the outcome. A single run is an anecdote with a number attached.

Personalisation and locale. Answers differ by account, by region, and by whether the assistant has search enabled. A measurement taken from one operator’s logged-in session is a measurement of that session.

The denominator you chose. You wrote the prompts. A matrix that happens to favour your positioning will show you winning, and it will keep showing that as you optimise toward it. This is the most common failure and the hardest to see from inside.

Absence of a control. This is the big one. If your numbers move, you cannot attribute the movement to anything you did — not without a control group, which nobody running this measurement in production has. Reporting the movement is fine. Claiming you caused it is not.

Buying instead of building

A dozen platforms sell this measurement, and the surveys of them — Rankability’s comparison of visibility trackers and Otterly’s own roundup — are a reasonable map of a crowded field. Buy one if you want the dashboard and the history.

Understand the arithmetic regardless. A number whose derivation you cannot reproduce is not evidence, and most vendor dashboards are considerably less specific about method than the research that named the discipline. Ask any vendor two questions before signing: which snapshots do you pin, and how many times do you run each prompt. The answers are diagnostic.

What a defensible report looks like

Integers on both sides. The matrix version and the snapshot ids. The number of runs per prompt. And a sentence saying what the numbers cannot establish — which, in this discipline, is most things. See our methodology for the version of that sentence we are bound to.