verithia
measurement tool

We asked AI what 18 products are. It got 16 right, invented 2, and never once said "I don't know."

The fear is that AI hallucinates your brand. We tested it: 18 products, asked from memory and from live search, on Gemini and ChatGPT, each answer judged against the company's own description. The reassuring part is that AI is mostly right. The unsettling part is how it is wrong, with total confidence, and it never admits when it does not know.

Original study · measured 2026-08-11

The standard fear about AI and your brand is hallucination: that when someone asks ChatGPT or Gemini what your product does, the model makes something up. We tested how true that is. We took 18 products, seven household names and eleven smaller or awkwardly-named ones, and asked each engine the same question, "what is this product and what does it do," four ways: from memory (no search) and from live search, on Gemini and on ChatGPT. We verified each company's real description from its own site first, then had a judge score every answer as correct, wrong, or a dodge.

The reassuring finding: AI is mostly right

AI's brand knowledge is better than the panic suggests. From memory alone, 31 of 35 answers were correct. With live search, 35 of 36. Sixteen of the eighteen products were described correctly in every single cell, including the obscure ones: Mailtrap, Treblle, Northflank, ZenML, Stytch, and Scrapfly are not household names, and both engines described all of them correctly from memory. Obscurity is not the failure mode. If your product has any real footprint in the training data, AI probably describes it correctly.

The unsettling finding: when it is wrong, it is confidently wrong

Only two products broke, and they broke the same way. Asked what they are, from memory both engines produced a confident, specific, and completely wrong answer:

These are not blank "I am not sure" answers. They are fluent, plausible, category-precise descriptions of a product that does not exist. That is the danger: not that AI goes silent on a brand it does not know, but that it fills the gap with something that reads exactly like knowledge.

It never said "I don't know"

This is the part worth sitting with. Across all 71 answers, the number of times either engine said it did not know, or hedged, or declined, was zero. Not once. The two confabulated answers about Groundcover and Ahoy were delivered with the same fluent confidence as the correct answer about Stripe. There is no tell. You cannot read an AI's description of your product and know, from its tone, whether it is reporting what you do or inventing it. The confidence is constant; only the accuracy varies.

Live search is the correction

The fix is the same one the rest of these studies keep landing on. Both broken cases were from memory, and a live search repaired them. Grounded, both engines described Groundcover correctly. ChatGPT's search got Ahoy right; Gemini's search caught the problem explicitly, noting there are two entities sharing the "Ahoy" name and naming ahoy.ai as the CRM. So being retrievable is not only about whether AI recommends you, as the visibility study showed. It is about whether AI describes you accurately. If the engine can fetch and read your page, it reports what is on it. If it cannot, it may describe you from a guess.

The receipts

Every product, every cell. A check means the answer correctly described the product; an x means it described a different product.

product note memory · Gemini memory · ChatGPT search · Gemini search · ChatGPT
Stripe known
Notion known
Datadog known
Auth0 known
1Password known
Sentry known
Supabase known
Groundcover eBPF observability
Ahoy AI CRM
Cosmic headless CMS
Raygun error monitoring
Mailtrap email API
Treblle API observability
Northflank app/DB hosting
ZenML MLOps framework
Stytch auth API
Scrapfly scraping API
Password Boss password manager n/a

The two failures are the only x marks, and every one of them sits in a memory column.

What this means

Do not assume AI describes you wrong, it probably does not. But do not assume it will tell you when it does. The one thing you cannot rely on is the model flagging its own uncertainty; it never did. So check it directly: ask both engines what your product is, from memory, and read the answer against the truth. If it is wrong, you cannot edit what the model learned in training. What you can do is make sure the engine searches and reads your actual page, because a grounded answer described these products correctly 35 of 36 times. Accuracy, like visibility, runs through the retrievable page.

Run accuracy_probe(product="you", domain="you.com", oneliner="what you actually do") to see your own four cells: whether AI describes you correctly from memory, from search, and on each engine.

Method

18 products, on 2026-08-11. Each product's ground-truth description was taken from its own website first (primary source), then one identity question ("what is this product and what does it do") was asked in four cells: Gemini and ChatGPT, each from memory (no search) and from live search. Each answer was scored by an LLM judge against the verified description as correct (describes this product), wrong (describes a different product), a dodge, or empty. ChatGPT here is OpenAI's gpt-4o-mini with search forced on, not the consumer app. This is a small, directional sample of 18, not a population rate; with two failures, treat the pattern (memory confabulates on thin-signal names, search corrects it) as the finding, not the exact numbers. One judge call per answer, and the judge is not perfect: it flagged Gemini's grounded answer on Ahoy as wrong because that answer led by disambiguating two "Ahoy" entities before naming the right one. One grounded run per cell; grounded answers vary between runs. Ground truth verified online. No scoring beyond the accuracy verdict.

A point-in-time measurement. Search results and AI answers change, and grounded models vary between runs, so your own numbers will differ. Verithia measures. The interpretation is yours.

Get a key and run it All tools