verithia
Fidelity check

We tested our own tool against two live AI surfaces. Here is exactly how close it is.

Most AI-visibility tools measure an API and call it "what AI says about you." Almost none tell you how far that is from what a person actually sees. We measured our own gap against two different live surfaces. Across 12 buyer queries, our API overlapped the logged-out Gemini app 58% and Google's live AI Overview 56%, and the surprising part is that both numbers are the same rate at which the engine agrees with itself between two runs.

Original study · measured 2026-08-12

Every AI-visibility tool has a dirty secret it does not print on the box: it usually does not measure what you see in the chat. It measures an API, a grounded model call, that resembles the consumer app but runs a different pipeline. The serious tools scrape the actual app; the cheaper ones call the API; and hardly any of them tell you which they did, or how close their number is to reality.

We think a company selling honest measurement should measure its own honesty first. So we ran 12 buyer queries through our API and, separately, through the logged-out Gemini app (no account, no memory, the clean baseline a stranger would see), and compared what each one cited.

The headline: 58% overlap with the live app

Across the 12 queries, the logged-out Gemini app cited 52 source domains. Of those, 30 (58%) also appeared in our API's citations across three runs, and 24 (46%) appeared in the ones our API cited consistently (in at least two of three runs). So on a plain reading: a little over half of what the real app cited, our tool also surfaced. Not a mirror. Not noise either.

Why 58% is the reassuring number, not the damning one

Here is the context that decides how to read that 58%. Grounded AI is not stable run to run. Our fan-out study measured that a single Gemini call captures only about 56% of the domains that three calls surface, because roughly half of what it cites is one-run churn. So:

The live app overlaps our tool at the same rate one of our own runs does. Statistically, the app looks like one more draw from the same distribution our API samples, not a different machine citing a different world. That is the real finding: at the level of the citation distribution, our API tracks the app about as well as the app tracks itself. The proxy is honest, as long as you read it as a distribution and not as a single verdict.

The catch: a single answer, from anyone, is unreliable

The flip side of that same fact: if half of what any single Gemini answer cites is churn, then no single answer, app or API, tells you whether you are cited. A one-shot "you appear in ChatGPT" is a coin flip dressed as a fact. This is true of our tool and every competitor's. The only trustworthy signal is what recurs across several runs. Any AI-visibility number that comes from one query, one time, is measuring noise.

Where the app and the API genuinely differ

It is not all noise; there is a mild, real lean. Look at what the app cited that our API never did: TechRadar, Tom's Guide, G2, Gartner, New Relic, Honeybadger. Established review sites and big-name brands. Our API, on the same queries, reached further into the long tail. Same model family, but the app's full search stack pulls harder toward mainstream, high-authority sources. And one category broke outright: for headless CMS, the app cited comparison blogs (Hygraph, dotCMS) while our API cited the products themselves (Sanity, Directus), zero overlap. Category-level divergence is real and you should expect it occasionally.

This lean has a documented cause, it is not just our observation. Google runs two different retrieval systems with different names. The consumer surface (AI Mode and AI Overviews) uses query fan-out, decomposing a prompt into a wide batch of sub-queries run across the search index, with independent studies putting the count at 8 to 12 per prompt. The API uses Grounding with Google Search, where the model fires its own, smaller set (we measured about seven). A wider fan-out means a larger candidate pool and, in Google's own words, "a broader and more diverse set of links," which is precisely the lean we observe. The same split is documented on the other side: developers report the ChatGPT app and the OpenAI API returning different web-search results for identical prompts, because the search path is implemented differently in the consumer app, which also cites sources from publisher licensing deals the API does not have. So the gap we measured is the expected result of two different pipelines, not a defect in either.

The receipts, query by query

category app cited also in our API app-cited that we missed
stripe alternatives 7 6 creem.io
best vector database 5 4 pingcap.com
best crm for startups 4 3 techradar.com
best vpn 4 3 tomsguide.com
best web scraping api 5 3 scrapebadger.com, aimultiple.com
best transactional email api 3 2 emailtooltester.com
best password manager 4 2 techradar.com, tomsguide.com
best authentication service 5 2 gartner.com, delinea.com
best error monitoring tool 6 2 honeybadger.io, new relic, glitchtip.com
best customer support software 2 2 (none)
best project management tool 3 1 g2.com, goodday.work
best headless cms 4 0 hygraph.com, dotcms.com
all 12 52 30 (58%)

A second surface, Google's AI Overview, lands in the same place

We did not stop at the Gemini app. We ran the same 12 queries against Google's live AI Overview, the AI answer most people actually see in Google Search, and compared its cited sources to our API. The result: 56% overlap on the union, 46% on the recurring set, statistically identical to the Gemini app's 58% / 46%. Two entirely different consumer surfaces, and our API tracks both at the same ~56 to 58%, right on the engine's own run-to-run variance. When two independent surfaces and the engine's self-consistency all land on the same number, that is not a coincidence. It is the noise floor of grounded AI, and our API sits inside it.

Three honest things that run surfaced:

category AIO sources also in our API what AIO cited that we missed
best vector database 4 4 (none)
best password manager 4 4 (none)
stripe alternatives 5 5 (none)
best crm for startups 10 6 HubSpot, Zoho, Zendesk, Zapier
best vpn 7 4 CNET, NordVPN
best authentication service 7 3 G2, Gartner, Rippling
best project management tool 9 2 Asana, Monday, G2, Forbes
best headless cms 8 2 Pantheon, Jamstack
8 queries where AIO showed 54 30 (56%) (4 queries showed no AIO)

What this means for reading any AI-visibility number, including ours

Two rules fall out of this, and they apply to every tool in this category:

  1. Trust the recurring set, never a single answer. Whether it is our tool or the live app, a citation that shows up once is a coin flip. Run it several times and act on what repeats. A vendor who reports a one-shot answer as "where AI cites you" is selling you noise.
  2. Know whether you are looking at the app or a proxy, and how close they are. Our API is a proxy, and we just told you the number: about 56 to 58% aligned with two live surfaces (the Gemini app and Google's AI Overview), which is within the engine's own run-to-run variance. No competitor we know of publishes their fidelity at all. We would rather show you the gap than hide it.

That is the honest position. Our tool is not a pixel-perfect copy of the Gemini app, nothing that calls an API is, but it samples the same distribution the app samples, and if you measure the way this study measures (several runs, the recurring set), it tells you the truth about where you stand.

Method

12 "best X" and category buyer queries, on 2026-08-11 (Gemini app) and 2026-08-12 (AI Overview). The app side is a single capture of the logged-out Gemini app per query (gemini.google.com, no account, no memory, so no personalization); the source domains are the citation attributions it displayed. The AI Overview side is a single capture of Google's live AI Overview per query (sources only); it appeared for 8 of the 12 queries. The API side is our tool run three times per query; "also in our API" counts consumer-surface domains that appeared in any of the three runs (the union). Domains are normalized to the registrable domain for cross-surface matching. The Gemini-app comparison is the same model family, which is the best case; the AI Overview comparison spans Google's full search stack, which fans out wider (independent studies put it at 8 to 12 sub-queries versus our API's ~7) and leans harder to authority sources. Limits worth stating: one consumer-surface capture per query (both surfaces are stochastic too, so more runs would likely raise the overlap), and the app's hidden "+N more" sources mean our count of its citations is a floor. The run-to-run baseline (56%) is our own measured figure from the fan-out study. No scoring, no verdict.

A point-in-time measurement. Search results and AI answers change, and grounded models vary between runs, so your own numbers will differ. Verithia measures. The interpretation is yours.

Get a key All tools