Methodology

How we measure, in full

“Measured honestly” is only a claim if you can check it. So here is the exact request we send to each answer engine: the endpoint, the model, whether search is on, which country and language we send, and how many samples we take.

We ask 50 buyer questions per site, written from your own website. 1 sample per question per engine each week, 3 in the monthly stability run. We show a direction in words only after 3 runs — below that, week-to-week movement is noise.

ChatGPT

via OpenAI
Endpoint
POST https://api.openai.com/v1/responses
Model
gpt-5.5
Grounding
On. The request sends tools: [{ type: "web_search" }], so the model searches the live web before answering.
Citations
We read the answer from the output_text content and keep every url_citation annotation attached to it as a cited source.
Country
Not sent. The Responses API request carries no location parameter, so the country you pick never reaches OpenAI — it only steers how the buyer questions are written.
Language
The language of your website. Your buyer questions are generated in that language and sent as-is.
Samples
1 per question per weekly run; 3 per question in the monthly stability run
Not the consumer app
The ChatGPT app answers a signed-in person with memory, custom instructions, chat history and an approximate location. Our call has none of that: a fresh, anonymous request every time. Read it as a clean baseline, not as one specific user's screen.

Perplexity

via Perplexity
Endpoint
POST https://api.perplexity.ai/chat/completions
Model
sonar-pro
Grounding
On. Search is built into the model — every answer is produced against live search results.
Citations
We take the citations array; when a response omits it we fall back to the url of every entry in search_results.
Country
Not sent. The Chat Completions request carries no location parameter — the market only shows up in how the buyer questions are worded.
Language
The language of your website, through the wording of the question itself.
Samples
1 per question per weekly run; 3 per question in the monthly stability run
Not the consumer app
Perplexity's own app personalises on account settings, search history and location. We always call the same model anonymously, so week-over-week differences come from the web, not from your profile.

Google AI Overviews

via DataForSEO (SERP API)
Endpoint
POST https://api.dataforseo.com/v3/serp/google/organic/live/advanced
Model
Google's live AI Overview for that search, requested with load_async_ai_overview: true
Grounding
Grounding is Google's own. We do not prompt a model here: we run the buyer question as a Google search and read the AI Overview that Google itself generates.
Citations
We read the items of type ai_overview and keep the url of every reference underneath them.
Country
Sent explicitly as location_code — 2840 for the United States, 2826 for the United Kingdom, 2528 for the Netherlands.
Language
Sent explicitly as language_code, taken from the language of your website.
Samples
1 per question per weekly run; 3 per question in the monthly stability run
Not the consumer app
A signed-in Google user gets results shaped by their account, device and precise location; we measure the country-level result. Not every search triggers an AI Overview — when Google shows none, that is a valid measurement with an empty answer, never an error.

A clean setup, every time

Every measurement is a fresh, anonymous request. What that rules out:

  • No personalisation — we are never signed in to a consumer account.
  • No memory, no chat history, no custom instructions.
  • No location history or device signals; only the country you configured, and only where the API accepts one.
  • No advertising or A/B variants we could influence.

Markets we can measure in: United States, United Kingdom, Ireland, Nederland, België, Deutschland, Österreich, Schweiz, France, España, Italia, Sverige, Danmark, Norge.

How the score is built

The AI visibility index runs from 0 to 100 and is nothing more than your average prominence, expressed as a percentage of the maximum:

index = (average prominence ÷ 3) × 100

  • The denominator is every measured answer, including the ones where you do not appear at all.
  • Sentiment is not part of the score. It is subjective, it is classified by a model, and it does not tell you what to do.
  • A question that turns out to be structurally unstable counts for half in both the numerator and the denominator, so noise cannot dominate your score.

“Not measured” is a real state

When an engine cannot be reached — a provider outage that survives our retries, for example — that answer is stored as not measured. It does not count as a zero and it does not quietly disappear: it drops out of both the numerator and the denominator, so your score describes only what we actually saw.

One case looks like a failure but is not. If Google shows no AI Overview for a question, that is a valid measurement with an empty answer: your business is not in it because there is no AI answer at all. We record it as measured with prominence 0.

Your first measurement is a baseline: one point, no direction. We will not show a delta or an arrow until there is something to compare against.

What we do not claim

An answer engine is not a search ranking: ask the same question twice and the wording moves. We measure a clean baseline, not one specific person's screen, and we never present a change as proof that our fix caused it — only as a change we observed after it.

See what AI says about your business.

A real measurement on 15 of your own buyer questions. No account, no card.

Scan my site free