Archived version, current as of 22 August 2026 — superseded. See the current disclosure page for GapWatch's present methodology and disclosure.
What we measure,
and what we don't
This page answers each item in the IAB AI-visibility disclosure framework, in the framework's own terms. The framework defines two tiers — Directional and Decision-Grade — and treats a programme under 50 queries as exploratory, below either. This page states which tier GapWatch meets on each criterion, including where it falls short. We would rather be checkable than flattering — a measurement vendor that discloses only its strengths has published marketing, not a disclosure.
Last updated 22 August 2026.
Query design
The queries are the instrument. A finding that cannot be reproduced from them is not worth reporting, so every panel publishes its full query set.
Raised from 26 on 22 August 2026. Each query carries a written rationale stating what it covers that no other query in the set covers; an automated check rejects a set with duplicate or missing rationales, because doubling a panel with near-duplicates satisfies a count and measures the same thing repeatedly.
Results are reported per segment and never blended. Evaluation-stage figures approach 100% by construction because the brand is named in the query; blending them with discovery destroys the discovery signal.
Transactional intent was added on 22 August 2026; before that date GapWatch covered three of the four. Distribution is published with every panel and results are segmented by intent type, which is what the Decision-Grade threshold requires.
GapWatch queries are designed against a client's actual products, competitors and stated customer problems. They are NOT grounded in real consumer search-volume data. The framework requires this distinction to be disclosed, and buyers should weigh it: a designed set can probe questions search-volume data never surfaces, and can equally miss what people actually ask.
Product-level and problem-level queries are generated from the client's actual products and stated customer problems, so subcategory coverage scales with what is sold rather than with a template.
Standing rule. An unanchored non-English query returns results for whichever market the provider defaults to, which is not a finding about the client's market.
Providers and execution
We record what actually answered, not the label we requested. The two diverge when a provider routes to a point release or a fallback, and that divergence is what a re-baselining decision turns on. Where a provider returns no version, it is recorded as unknown rather than defaulted to the requested label.
All are queried with web search or grounding enabled, via official APIs — active query simulation, not passive panel observation or platform-native data. Results are reported per provider; a blended cross-provider figure describes none of them.
Generative answers vary between identical prompts. One run is a sample of one.
Every percentage is computed against observations actually captured for its own segment. Using scheduled observations lets a provider outage deflate every entity equally, turning a visibility measure into a measure of provider uptime.
Where a provider returns fewer observations than scheduled, the shortfall and its cause are published alongside the results.
Prominence
Beyond ordinal rank: whether the brand was the only tracked entity named (a stronger position than leading a list), how many times it was referenced, and how far into the answer the first mention falls — a brand named once in a closing sentence is weaker than one named in the opening, and rank alone cannot see that.
GapWatch does not render responses, process screenshots, or measure above-the-fold visibility — querying via API means no rendered surface exists to measure. Rendered position is therefore not reported rather than approximated from text. This is a structural limitation, not a roadmap item.
Variability
The framework's example of good practice is “22%, ±4 points”. As of 22 August 2026 GapWatch reports ranges rather than point estimates.
Measured 22 August 2026 from two runs of an identical query set, 15 and 27 July. Perplexity moved 10% to 20% between identical runs — the clearest reason point estimates were the wrong format. Caveats published in full: the runs were 12 days apart rather than within 7, the set was 10 queries rather than the 56-query standard, and a 0-point band on our own 0% discovery figure is a floor effect that must not be read as stability. House default where no slice-specific band exists: ±5 points, the worst observed rather than the median.
Same queries, same providers, same run count, same location, same language. A comparison across any changed parameter measures the change rather than the variability, and is refused rather than warned about.
No sampling distribution is assumed and no standard error is computed. A band states how far this instrument moved against itself between two runs — the honest floor on how precisely a single figure should be read. Calling it a confidence interval would import guarantees the design does not earn.
Variability differs by both. A discovery figure from a provider that moves 9 points is not the same quality of number as an evaluation figure from one that moves 1, and a single blended band would hide precisely that.
Where we fall short
Two criteria where GapWatch meets Directional but not Decision-Grade. Both are stated here rather than in a footnote.
Testing cadence — meets Directional, not Decision-Grade
The framework sets Directional at monthly or quarterly collection and Decision-Grade at weekly or more frequent. GapWatch panels are run on request and on freeze dates. A GapWatch figure describes the date it was captured. Where a client needs trend data we run repeated panels on a fixed cadence and report the series — but the standard product is a dated snapshot with a variability band, not weekly monitoring.
We do not describe our standard panels as tracking, monitoring, or continuous.
No Google AI Overviews coverage
GapWatch measures four conversational assistants. It does not measure Google AI Overviews, which for many categories is the highest-traffic AI answer surface in existence.
There is no public API for AI Overviews; reaching it requires a third-party SERP vendor. As of 22 August 2026 we have not contracted one, because the leading vendor is a defendant in active litigation brought by Google over reselling extracted search content for a fee — which is what incorporating it into a paid client report would be. We are not willing to place that risk into a client deliverable, or to claim a surface we have not legally cleared.
This matters against the framework's Platform Coverage threshold, which asks for platforms representing a substantial majority of consumer AI traffic. The framework itself notes AI Overviews reach over 2.5 billion monthly users and appear on almost half of searches. Any GapWatch finding should be read as being about conversational assistants specifically, not about “AI search” as a whole.
Other surfaces we do not measure
Absence from this list is not a judgement about a surface's importance. It means GapWatch has no measurement of it and makes no claim about it.
Panels are frozen on a date and audited before release. Numbers already shipped in a client document are never silently changed by a later re-run.
Claude is also one of the four measured providers. This conflict is disclosed on every artifact where it applies rather than only here.
Corrections to this page are welcome and are published with a date rather than applied silently. See also methodology and limitations.