← All insights

Mentioned Isn't Cited: What 1,620 AI Answers Reveal About Brand Visibility in DACH

Discuss with AIClaude ↗ChatGPT ↗

When someone recommends a company to you, two separate things happen. They say the name. And they give you the address so you can actually walk in. In answers from ChatGPT, Gemini and the rest, those are two distinct events, and they drift apart further than most people assume. The name in the running text is the recommendation. The clickable link to your domain is the address. A brand that gets praised but never linked exists in the conversation and vanishes at the door.

That gap is what I measured on 27 July 2026. 180 commercial search scenarios, six industries, three markets (Germany, Austria, Switzerland), three AI engines, three runs per combination: 1,620 controlled API queries, all raw data open. The result is not a ranking list. It is proof that two metrics people constantly conflate often point in opposite directions.

Two numbers that don’t line up

The hiking-boot maker LOWA is named in 92.2% of the tested answers in Germany. lowa.com is linked in 26.7%. So two of every three recommendations end without an address. Run the self-storage industry through the same analysis and the picture inverts completely: the brand is named in the text only 7.6% of the time on average, but linked in 48.1%. Here the address is almost never missing – the name often is.

That is the central observation of this baseline: mention and link diverge by industry, sometimes one way, sometimes the other. Treat visibility as nothing more than “do we show up in AI answers?” and, depending on the industry, you are measuring the wrong half.

FIG 01 · MENTION → CITATION BY INDUSTRY – THE DECOUPLING 0 20 40 60 MENTION % CITATION % ECOM 56.4 SAAS 45.9 YMYL 26.2 IND 25.4 TRAVEL 15.0 LOCAL 7.6 17.3 LOCAL 48.1
Six industries, two metrics. E-commerce and SaaS lose ground from mention to link; Local flips it (blue): barely named, yet nearly every second answer links. OWN ANALYSIS · AS OF 07/2026

Defined briefly, with everyday examples

Before this goes deeper, three terms without which the rest stays fuzzy. I use them throughout with exactly this meaning.

The mention rate is the share of answers in which your brand name appears in the running text. If someone asks about hiking boots and the answer writes “Lowa, Meindl and Hanwag are established brands,” that counts as a mention for all three. Whether a link to each site sits alongside it does not matter for this number.

The citation rate is the share of answers in which your domain shows up as a clickable source below or inside the answer. That is the address. A brand can be highly cited and rarely named (the Local case) or often named and rarely cited (the e-commerce case). Both metrics are averages across 90 runs per provider, measured independently of one another.

The grounding rate, finally, tells you how often a model actually looks things up live on the web instead of answering from memory. A model that does not ground can know and name you from its training, but cannot link a fresh source. A mention without evidence I call a hallucinated mention – the model knows you exist but simply did not look just now. Its counterpart is the attributed citation: a mention that sits on a source actually retrieved and carries a link. The gap between the two is not a statistical nuisance. It is the real subject.

Why a model praises without linking

The answer sits in how the engines behave, and LOWA shows it in miniature. Split by model, the German e-commerce measurement looks like this: GPT-5 Mini names LOWA in 86.7% of runs and links in 40.0%. Gemini 3.5 Flash names it in 93.3% but links only in 26.7%. Gemini 3.5 Flash-Lite pushes it to the extreme: 96.7% mention, 13.3% citation. The harder a model leans on its built-in brand knowledge, the more often the name falls on its own.

FIG 02 · TWO WORKING STEPS IN ONE ANSWER EXPLAINING PART EVIDENCING PART Parametric brand knowledge Synthetic comparison list MENTION in text no link Retrieved pages (grounding) Defensible single claim CITATION with link the evidence wins
The same answer names from memory and evidences from retrieval. LOWA wins the left step almost always (92% named), the right one often not (27% linked). OWN ANALYSIS · AS OF 07/2026

This comes down to two working steps that meet in the same sentence (FIG 02). The explaining part the model builds from its brand knowledge: LOWA, Meindl and Hanwag belong to common knowledge about hiking boots and appear almost reflexively. The evidencing part it pulls from the retrieved pages, where in the brand’s place a retailer, a review portal or a guide often supplies the citable claim. The brand gets praised; the evidence gets linked.

For you that means a high mention rate with a low citation rate is not an awareness problem. You are known well enough to be named. You are simply not the page the model presents as its source. It is the same gap I described elsewhere as retrieved but not cited, just one stage earlier here: named, but not evidenced.

The local paradox, resolved honestly

That leaves the opposite direction. Why do self-storage providers of all people collect nearly every second link while their names barely appear in the text? Two mechanisms work together, and one of them is a measurement artifact I have to disclose, or the number overstates the case.

The first mechanism is real and the more interesting one. Local questions force the model to ground. “Which self-storage providers with easily reachable locations are there in Berlin?” cannot be answered sensibly from memory; the model has to fetch current, location-specific data. And that sits in directories, in map services, on the providers’ own location pages. That is exactly where the link then goes. Where a question would be wrong without fresh grounding, the engine reliably links the grounding source. This matches what I described about the role of Maps and reviews in grounding.

The second mechanism is the measurement error, and it concerns the low mention rate. My mention metric checks the exact provider name including the city suffix, so “MyPlace Wien.” The answers, though, say “MyPlace.” The model does name the brand – just without the city – and the exact match does not count that as a hit. In the raw data it is unambiguous: “MyPlace” appears in the text, “MyPlace Wien” never, myplace.at gets linked. So the real mention rate is higher than 7.6%. The paradox shrinks as a result, but it does not disappear: the link advantage holds, because local answers ground structurally.

B2B industry and SaaS drop on the links for the opposite reason. Their questions can often be answered from brand knowledge, without a fresh source being needed. The model names Salesforce because it knows Salesforce, not because it just read salesforce.com. Mention yes, evidence no.

Industry explains part of the spread; the intent behind the question explains more. Each of the 60 scenario types carries a buying stage, from first discovery to transaction. Count the retrieved sources per run and the prompt type shifts citation density markedly (FIG 03): from 18.8 citations for evaluation questions down to 11.5 for shortlists.

FIG 03 · CITATIONS PER RUN BY BUYING STAGE 0 5 10 15 Cites/run Evaluation18.8 Comparison18.7 Validation17.1 Alternatives16.5 Task-solving13.7 Customer support13.6 Discovery13.2 Transaction13.1 Shortlist11.5
The engine evidences most densely where options are weighed (Evaluation, blue), not where the purchase happens (Transaction, Shortlist). To get cited, you have to serve your category's comparison questions. OWN ANALYSIS · AS OF 07/2026

This runs against everyday intuition. You would expect a purchase-near question (“Where in Vienna can I reserve a storage unit online at short notice?”) to produce the densest sourcing. In fact the model pulls the most evidence where options are weighed, not where the purchase happens. An open comparison forces it to evidence several options against one another. A transaction question gets a shorter, more decisive answer with fewer sources. Concretely, these three prompts sit in the same industry and draw completely different behavior:

  • Discovery: “Which hiking-boot brands and online shops are relevant for day and multi-day tours in Germany?”
  • Task-solving: “Which waterproof hiking boots suit wide feet and wet, rocky Alpine trails?”
  • Comparison: “Compare the three most suitable CRM systems for a 50-person B2B team by rollout effort, automation, reporting, data export and total cost.”

The comparison question produces the densest, best-linked answer. The discovery question names many brands and links few. In practice this means: if you want to be cited, you have to serve the comparison and evaluation questions of your category, not just the brand or purchase intent. That is where the engine hands out the links.

A second effect sits underneath and concerns only GPT-5 Mini, because that model is allowed to skip the search. On comparison questions it grounds in 98% of runs, on evaluations in 94%. On shortlists the grounding rate falls to 57%, on customer support to 46%. Where the model believes it already knows the answer, it searches less often and links accordingly less. So intent steers not only how many sources get pulled, but whether a search happens at all.

Which pages the models cite

When the link does get handed out, where does it point? I counted the 13,059 real citation URLs from GPT-5 Mini by page type. The result is clear: roughly three of four citations lead to deep, specific subpages, not to homepages. Only a scant quarter lands on a homepage or a top level. Category and product pages from shops make up around 7%, blog and guide content a good 5%, explicit test and comparison pages a good 2%.

The most revealing block is the third-party sources. Official registers and trade directories show up prominently as cited sources, especially in mortgage advice: ris.bka.gv.at, finma.ch, wko.at, vermittlerregister.info. In the Swiss SaaS analysis, alongside the provider sites stand the data protection authority edoeb.admin.ch and the Microsoft documentation learn.microsoft.com. So the model evidences its claims not only with the brand, but with the authority that confirms the brand.

From this follows an uncomfortable consequence for homepage optimization. The engine wants the page that answers a concrete question concretely, not the company’s calling card. A well-kept location page, a clean spec sheet, a guide page with a verifiable recommendation get cited sooner than the homepage. Polish only the homepage and you optimize the very spot that gets the link least often.

The zeros say the most

Five providers sit at 0.0% on both metrics. These outliers teach more than any top position, because they expose different technical causes.

The clearest case is weclapp in Switzerland. Across all 90 Swiss SaaS runs the brand is not named once, not linked once. This is not a phrasing problem but a missing entity assignment: for the models, weclapp in the context “CRM Switzerland” is simply not a relevant entity. The slot you would expect it to hold goes to bexio, the local incumbent, with 69 citations in the same analysis. Where a strong local entity holds the slot, a brand weaker anchored in the market does not even make it into the mention. For comparison: in Germany weclapp is at least named 15.6% of the time and cited 21.1%. Same software, a different market, a different entity status.

Mammut in the German hiking-boot test is a different zero type: 6.7% mention, 0.0% citation. The brand is well known and named here and there, but in the specific hiking-boot segment never presented as a source. That is a category-fit problem, not obscurity. Mammut stands for mountain sport broadly; for hiking boots specifically the specialist brands win the evidence.

The remaining zeros (Lagerplatz.de in Berlin, Boxroom in Vienna, FW Finanzarchitektur in Austria) follow the same pattern: weak or missing anchoring in exactly the directories, registers and map data from which the engine draws its evidence for that question. Whoever does not appear where the model grounds does not exist for the model. Visibility does not begin with content, but with whether the grounding data layer knows you at all.

What I advise a client with a number like that

The most common case in consulting is the high mention with the low citation, the LOWA type. The first reflex, to produce more brand content, misses the problem. You are already being named. What is missing is presence at the spot where the engine fetches its evidence, and that often does not sit on your own domain.

Hence the concrete lever: build presence in the third-party sources the model grounds from. For local and for regulated businesses those are trade directories, official registers and map data; for products the independent test and comparison sites; for B2B the documentation and integration ecosystems. Your own domain stays important, but it does not suffice as the sole evidence source when the engine structurally hands out its citations elsewhere. The goal is no longer reach, but citability: to appear at the three or four places the model actually retrieves for your category.

In practice that means three steps in this order. First tighten the entity – register entries, directory profiles and map data with clean, consistent information – otherwise it stays at the zero value. Then make the one page per central question citable, with a verifiable claim high up rather than at the end of the text. Only after that scale content. Reversing this order is the most expensive mistake, because scaled content without an entity foundation produces exactly the 0.0% that weclapp shows in Switzerland.

The API surface is measured, not the web chat

One question decides the reading of any AI visibility number: which surface was measured? This benchmark queries the models directly through their programming interfaces, routed via OpenRouter, for Gemini onward to Google’s Vertex search, for GPT-5 Mini to the virtualized Exa search. That is not the same as chatgpt.com, gemini.google.com or perplexity.ai in the browser. They are two different environments, and their numbers are allowed to diverge.

For an open benchmark the interface is the cleaner choice, for one reason: reproducibility. Fixed model versions, a fixed tool contract, one machine-readable piece of evidence per run, and an open panel. Anyone can take the published panel_v1.json and rerun the same queries. A web surface does not allow that. It is a black box that runs A/B tests continuously, alters system prompts silently and personalizes by location, cookies and conversation history. So the interface measures visibility in API integrations, AI agents and enterprise applications; the web surface measures what a single human sees in the browser. Both are legitimate. They are just different questions.

The price of this choice belongs on the table. On the interface, GPT-5 Mini’s web search is an optional tool; the model skipped it in 22.6% of runs – whereas on chatgpt.com the search is usually forced for recommendation questions. So GPT-5’s real citation rate in the web chat is probably higher than measured here. Personalization is missing too. From this follows the one rule I apply to every GEO number: never compare values from different surfaces without naming the surface. For daily monitoring of what your client sees in the app, a service like peec.ai fits better, since it queries exactly these web surfaces. For a reproducible model benchmark, the API surface is the more honest basis. A later wave can place both side by side.

When this holds and when it doesn’t

These numbers are a baseline from 27 July 2026, not a trend. They rest on three engines (GPT-5 Mini plus Gemini 3.5 Flash and Flash-Lite), on brand-open questions and on a single measurement snapshot. For navigational or pure brand prompts other patterns apply, which I deliberately excluded here.

Two methodological limits are decisive for the reading. The mention metric matches the exact name including the city suffix and thereby undercounts local brands, as the MyPlace case shows. And the content-format analysis rests on GPT-5 Mini alone, because Gemini delivers its sources through masked Vertex redirects whose target path I cannot classify without further resolution. Both limitations do not weaken the core finding; they only bound how far individual percentages can be pressed. The raw data lies fully open – anyone who wants to recompute can do so against panel_v1.json, wave_1_results.json and wave_1_summary.json.

Conclusion and a test for your own brand

Mention and citation are two metrics, not one. Depending on industry and question intent they point in different directions, and whoever measures only one is managing the wrong half of their AI visibility. The good news sits in the split itself: a high mention with a low citation is a solvable problem, because the awareness is already there. Only the address is missing.

To take with you, a short self-test. Take your most important category question and put it to ChatGPT, Gemini and Perplexity, three times each:

  1. Is your brand named in the text? (mention)
  2. Is there a clickable link to your domain? (citation)
  3. If linked: is it your homepage or a deep, answering subpage?
  4. Which third-party sources are cited beside or instead of you? (directory, register, test, competitor?)
  5. Does the picture change when you switch from a comparison to a purchase question?

The answers tell you which of the described cases you are in: named and linked, named but unevidenced, or not even in the grounding data layer. Each case has a different lever, and none of them is a blanket “more content.”

FAQ

What is the difference between mention and citation in AI answers? The mention is the naming of your brand in the running text. The citation is the clickable link to your domain as a source. The two are handed out independently. In this baseline e-commerce is named 56.4% and linked 17.3%; Local is named 7.6% and linked 48.1%.

Why is my brand named but not linked? Because the model knows your name from its training and names it in synthetic comparison lists, but hands the evidence link to the page that supplies a concrete, retrieved claim. That is often a retailer, a review portal or a register, not your own domain.

Why do local providers have such high citation rates? Local questions cannot be answered from memory; the model has to fetch current location data from directories and map services and reliably links these grounding sources. Part of the low local mention rate is also a measurement artifact, because the exact name match requires the city suffix.

Which pages should I optimize for citations? Not the homepage. Roughly three quarters of citations lead to deep subpages: answering guides, category and product pages, spec sheets. The second lever is the third-party sources the engine grounds from, such as trade directories and official registers.

Why query by API and not through the web chat or a monitoring tool? Because an open benchmark has to be reproducible. The API query via OpenRouter fixes the model version and the tool contract and returns one machine-readable piece of evidence per run, so anyone can rerun the measurement. Web surfaces like chatgpt.com personalize and change continuously. These numbers hold for the API surface; what a user sees in the browser is captured better by services like peec.ai. Values from different surfaces should never be compared without naming the surface.

Which engines and data underlie this? GPT-5 Mini plus Gemini 3.5 Flash and Flash-Lite, measured on 27 July 2026 across 1,620 runs. All raw data, prompts and providers are published as open data on the research page.

Sources and status

Own study “DACH AI Source Benchmark 2026,” wave 1, collected on 27 July 2026.

How it was measured. Each of the 180 brand-open questions ran three times against three search-backed engines, so 1,620 queries. The surface was the OpenRouter chat completions interface with web search enabled (openrouter:web_search), at most five search calls per answer, sampling left at the provider defaults (temperature not pinned). The three engines are GPT-5 Mini, Gemini 3.5 Flash and Gemini 3.5 Flash-Lite. The third slot was originally Perplexity Sonar; the model returned an HTTP 404 over the endpoint and did not fulfill the shared tool contract, hence the switch to Flash-Lite. That leaves two of three engines at Google, so provider diversity is limited. A run counts as a valid observation only if at least one search call took place and at least one machine-readable HTTPS citation was present; the raw answer along with the model returned was stored. An answer without a search or without a citation counts as a grounding failure, not a zero value. The mention is the naming in the text; the citation is a clickable source shown in the answer text. The name match checks the exact provider name including the city suffix and thereby undercounts local brands, as the MyPlace case shows.

Why these brands. The selection was fixed before the first measurement and never changed afterward. Per industry and market there are five comparable providers, 90 provider rows in total, all verified before the run. A provider entered the panel only if an official source evidenced the category membership, a current sourcing or distribution route existed in the respective market, it compared sensibly with the other four in the cell, and it was not added or removed in reaction to results. The brand names are not part of the questions; the questions were brand-open, so the engine itself decides whom it names. The results were not used to rewrite the question population afterward. Alongside the provider check ran an independent second review of taxonomy, quotas and phrasing, a bias review, and for mortgage advice a separate YMYL check with register, ownership and compensation evidence.

No client, no mixing. For 27 July 2026 it is declared that in none of the six industries (hiking boots, CRM, industrial marking systems, self-storage, mortgage advice, family hotels) does a current or former client relationship exist. No tested provider is a client, no brand was chosen to support a desired result. This open panel stays strictly separate from client data and case studies; no client-derived sample enters it. Should a relevant business relationship arise in future, the declaration will be re-checked before the next wave.

Limits. This is a baseline, a single snapshot, not a trend. It rests on three engines and brand-open questions; for navigational or brand prompts other patterns apply. The content-format analysis rests on 13,059 real GPT-5 Mini citation URLs, because Gemini delivers its sources through masked Vertex redirects. GPT-5 Mini skipped the search in 22.6% of runs.

Aggregate tables, methodology and the full open-data files (panel_v1.json, wave_1_results.json, wave_1_summary.json) lie open at https://eullrich.com/en/research/quellenbenchmark-2026/.