Methodology and disclosure

How GapWatch measures

Version 1.1 · 21 September 2026

In August 2026 the Interactive Advertising Bureau published Measuring Visibility in the AI Era, the first industry guidance for measuring how brands appear in AI answers. It sets out two quality tiers — “directional” and “decision-grade” — across nine criteria, together with a set of disclosures it asks every measurement provider to make.

We publish our position against every one of them. Where our measurement meets the guidance, we say so. Where we are extending it, we describe what we are adding.

How our measurement works

We put to leading AI assistants the questions a customer would ask — about a category, a market, a brand and its competitors — and record every answer: which brands are named, how each is described, and which sources the answer rests on.

Four principles run through all of it.

Questions are fixed in writing before collection begins. Nothing is adjusted after we see the answers. That is what makes a second measurement a genuine comparison rather than a new opinion.

Every question is asked repeatedly. An AI answer is a distribution, not a single value, so we ask each question several times of each assistant and report the range.

Every figure carries its denominator. No percentage appears without the number it was calculated from.

A person checks what decides a finding. Automated classification is reviewed by trained human readers, and their level of agreement is measured and reported.

The nine criteria

Sample size

Each question is asked three times of each assistant. The number of responses is disclosed in every report, alongside a documented process for assessing consistency between runs. We are studying how stability changes with additional repetitions, so that the number of runs in each programme rests on measured evidence rather than convention.

Query volume

Programmes intended to inform decisions are built on question sets above fifty questions — the guidance's threshold for characterizing a category — with total volume, distribution by category and variety of question format all disclosed. A representative national programme uses fifty-nine questions across six market and language cells. Smaller exploratory scans are labelled as such.

Prompt type coverage

The guidance distinguishes four intent types: informational, comparison, recommendation and transactional. Our question sets cover all four, and results are reported segmented by type. That segmentation can show a brand performing very differently depending on whether it is already named in the question — which in our current work has been among the most revealing findings.

Extending: we are expanding transactional questions, the moment closest to purchase. We build these with each client, drawing on their knowledge of how customers actually order and, where the client makes it available, on the real search queries that reach them.

Testing cadence

Baseline measurements establish a dated reference point. Ongoing programmes are specified at a seven-day or more frequent cadence, which is what the guidance requires for decision-grade measurement.

Reproducibility

We measure how much identical collections vary, and report results as a range with a stated confidence level rather than as a single figure. Our process for verifying consistency is documented. Programmes are designed with a seven-day repeat on the identical question set, so that week-over-week movement can be separated from the measurement's own natural variation; the first such repeat on a current programme is under way.

Programmes are designed to include fixed-answer control questions on every pass. If those move, the assistants have changed — which tells us whether a shift reflects the brand or the instrument.

Data validation

Every brand is disambiguated before anything is counted. Brand names frequently collide with ordinary words, places and people, and a measurement that does not separate them is measuring something else. We report the share of collected material excluded for this reason, which for brands whose names are also common words can exceed half of everything collected.

Classifications are validated by trained human readers. On our sentiment instrument, two independent readers agree at 0.84 (Krippendorff's alpha) on a deliberately difficult sample, and our automated classification is measured against that standard. Where automated classifiers disagree, a person resolves the case.

Extending: where clients authorize access to their own platform data — Google Search Console and web analytics — we validate our measures against it directly. These are the platform-native and first-party signals the guidance names.

Methodology documentation

Question sets, markets, languages, competitive sets and classification rules are written down and frozen before collection, and versioned whenever they change. We document how raw answers become numbers and how each assistant is handled individually. Model identifiers are recorded for every answer and are available on request.

Platform coverage

Our programmatic measurement covers four leading AI assistants, which together account for approximately 91% of web visits to standalone AI chatbots (Similarweb, August 2026). That figure describes standalone assistants, and does not include answers generated within search results.

Extending: Google AI Overviews appear within search rather than as a standalone assistant, and Google offers no compliant programmatic route to observe them. We are adding coverage by direct observation — a person running the questions and recording each answer and its sources — and, where a client authorizes it, will measure through their own Search Console, which shows how their brand performs in AI Overviews directly.

Multi-platform aggregation

Results are reported separately for each assistant, with the variation between them shown. We do not combine assistants into a single visibility score: they often disagree in ways that are informative in themselves, and a composite would conceal that.

How we obtain our data

The guidance asks every provider to state its access method and whether that method complies with the terms of the platforms it measures.

  • AI assistants are queried through their official APIs.
  • Public conversation is drawn from public content through third-party data providers, each identified on request.
  • Surfaces with no compliant programmatic route: our approach is direct observation by a person, as any user would see them.
  • Client-authorized data — a brand's own Search Console, analytics or community spaces — can be included with the client's permission.

We read each platform's terms before building any collection route. We do not use fingerprint evasion, stealth browsing or persona accounts, and where a platform offers no compliant route we observe it directly rather than working around its controls.

What we cover

Coverage is not a fixed list of platforms. It is every surface legitimately open to us, plus any surface a client chooses to open. Scope is agreed with each client before measurement begins, and for every surface we state the method, the cadence and the limits.

What our measurement is for

Our measurement shows what AI assistants say about a brand, what those answers rest on, and how that changes over time. It is designed to inform strategy and to establish whether work moved the signals it was meant to move.

We measure; we do not sell optimization services. The people who evaluate a result are not the people who produced it, which is what makes the evaluation worth having.

We do not claim to measure sales, persuasion or causality, and we describe our measurement as “decision-grade” only where a written assessment against every criterion supports it.

Terms in quotation marks are as defined in the IAB's Measuring Visibility in the AI Era, published 3 August 2026. This page is versioned; material changes are dated and recorded.

Previous version: 22 August 2026. See also methodology and limitations.