AI & Automation

AI Will Cite You Anywhere. It Only Recommends You at Home: 283,215 Citations Show Brand Mentions Halve Outside Your Category

2026.08.11 · 16 views
AI Will Cite You Anywhere. It Only Recommends You at Home: 283,215 Citations Show Brand Mentions Halve Outside Your Category

On August 5, 2026, Kevin Indig published a study simultaneously on Search Engine Land and his own Growth Memo, built on Semrush AI Visibility Toolkit data from US ChatGPT: 1,094 categories, five prompt variants per category, January through June 2026, totalling 283,215 citation observations and 76,493 named-brand-mention observations. The finding fits in one sentence. Step outside your core expertise and AI cites you as a source 50% of the time but names your brand only 25% of the time. Stay inside it and those numbers are 74% and 44%. Being cited and being recommended were never the same event, yet nearly every AI visibility dashboard adds them together into a single score. This became measurable only because the tooling finally grew up. Over the past 18 months, AI visibility tracking turned from a buzzword into a capital market: Profound closed a $96 million Series C at a $1 billion valuation on February 24, 2026, taking total funding past $155 million; Semrush prices its AI Visibility Toolkit at $99 per month for 25 prompts; Otterly.AI starts at $29 per month. Every one of them sells the same object — a percentage called "AI visibility." The trouble is what sits in the numerator. Similarweb, tracking US desktop users, found that people who saw a brand recommended by ChatGPT were 2.5 times more likely to visit that brand within seven days than a competitor, and those visitors viewed 12 pages and stayed 11.8 minutes versus 6.5 pages and 5.6 minutes for everyone else. Revenue follows the mention, not the citation. For Taiwanese SMBs and contract studios, this study is not a new metric. It is a subtraction: narrow your content strategy instead of widening it. Below are the full numbers, what three kinds of readers should do this week, whether any tool actually separates the two signals, and how to measure it yourself without a subscription.

Share:

AI Will Cite You Anywhere. It Only Recommends You at Home.

At 11:00 am on August 5, 2026, Search Engine Land published Kevin Indig's Does topical focus make your brand more visible?. He ran three tests against a US ChatGPT dataset supplied by Semrush. The sharpest numbers sit in the first one. When a brand appears in a category semantically distant from its core expertise, 50% of those appearances are citations (your URL listed as a source), only 25% are mentions (your brand name written into the answer), and just 9% are both. In categories close to its expertise those three numbers become 74%, 44% and 34%. Same brand, same window — change the topic and the odds of being recommended roughly halve.

This is measurable now because the measurement infrastructure just arrived. Over the past 12 to 18 months AI visibility went from a slide-deck noun to a real software category, and the defining event was Profound announcing a $96 million Series C at a $1 billion valuation on February 24, 2026, led by Lightspeed, pushing total funding past $155 million, with the company claiming more than 700 enterprise customers covering over 10% of the Fortune 500. Only a sampling database at that scale makes a six-month panel regression across a thousand categories possible. This research is a by-product of the market maturing.

The gap between tools is not data volume. It is whether they are willing to separate the two signals. Semrush AI Visibility Toolkit runs $99 per month for one domain and 25 prompts and supplied the data here. Peec AI starts around $99 per month and runs about $212 at Pro. Otterly.AI is $29 per month at Lite and $189 at Standard. Profound starts near $499 per month. That is a 17x price spread, and almost every one of them puts a single "Visibility Score" on the front page, collapsing citations and mentions into one number. The category is expanding, not consolidating — and the usual side effect of an expansion phase is that everyone races to define the metric and nobody races to split it.

For a Taiwanese SMB or contract studio, the practical takeaway is not to buy another tool. It is that if you are currently "adding topics to widen the funnel," this dataset says the thing you are buying is citations, not customers. What follows: the full numbers, this week's actions for three reader types, the only line in the tool comparison that matters, and how to run the same measurement yourself for free.

The Details and the Full Numbers

Start with scale, because it decides whether the conclusion is worth acting on. The study uses Semrush AI Visibility Toolkit US ChatGPT data covering 1,094 categories, five prompt variants per category per month, January through June 2026. The breadth models use 283,215 citation and 76,493 named-brand-mention domain-category observations, each paired with the following month's outcome. The relatedness analysis uses 45,578 expansion appearances by 1,458 mapped brand entities. The author states plainly in the methodology that these are associations after controls, not causal effects.

Three results are worth writing down:

  • Citations ignore the topic. Mentions do not. Looking only at citation-without-mention presence, distant categories score 41% and close categories 40% — effectively identical. If your page meets the ordinary criteria for being a source, AI will cite you on almost any topic. Getting it to say your name requires the topic to sit in your core.
  • Breadth is not punished. Shallowness is. On the citation side, a brand hitting only 1 of a category's 5 prompt variants scores +0.012, rising to +0.062 at 5 of 5. No spread-thin penalty. Mentions run the other way: 1 of 5 is associated with -0.051, and only at 5 of 5 does it turn slightly positive. "Touch everything a little" and "fully own everything you touch" are different strategies, and the first one damages brand mentions.
  • Industries differ. In finance the citation-breadth association climbs from +0.054 to +0.139; real estate goes from +0.021 to +0.135. Category expansion is viable there. In legal, the mention association stays negative even at full 5-prompt coverage at -0.058, and healthcare sits at -0.037. On high-stakes topics, being a credible source is not the same as being the recommendation.

A second dataset explains why mentions are worth more. Search Engine Land's June coverage of the Similarweb study found users were 2.5 times more likely to visit an AI-recommended brand than a direct competitor within seven days. In travel, 12% of users visited Kayak after a Kayak recommendation versus 3.4% who went to Skyscanner. Those visitors viewed 12 pages and spent 11.8 minutes on site against 6.5 pages and 5.6 minutes for non-AI-influenced traffic. Most importantly, 55.9% of AI-influenced visits arrived through search — in your GA4 they look like organic search, not like an AI recommendation.

What Three Kinds of Readers Should Do Now

Brand owners and small business operators

  • Write the category you intend to own as one sentence. Limit: one. A firm doing web design plus SEO plus LINE marketing plus brand consulting reads to an AI as four shallow categories, not one full-service vendor.
  • Audit this quarter's content calendar. How many new topics did you open? If it is more than one, stop.
  • If you run a clinic, law firm, accounting practice or aesthetics business, note that legal and healthcare are the only two industries in this study where expansion stays negative. Move budget into depth on a single procedure or practice area, plus third-party reputation.

Marketers and SEO practitioners

  • Build a five-question depth sheet for your core category: five ways a real customer would ask, run once a month on a fixed date against ChatGPT, Gemini and Perplexity, logging two independent columns — did my URL appear (citation), did my brand name appear (mention).
  • Stop reporting a single visibility score upward. Report three columns instead: citation rate, mention rate, and both-together rate. They frequently move in different directions.
  • Drop "number of new topics" as a content KPI. Use depth: of the five phrasings in your core category, how many name you. Five of five means you own it. One of five is negative in the data.

Developers and contract studios

  • Turn the 5-questions-by-3-models monthly snapshot into a scheduled script writing to a database. Minimum fields: date, model, prompt, cited, mentioned, and the competitor names that appeared in the answer. This is the table client dashboards are actually missing.
  • Wire up the Search Generative AI performance report (shipped June 3, 2026) as the first-party control and place it beside your sampled data, so clients can see which column is real impressions and which is an estimate.
  • Run a category-consolidation audit: cluster existing articles semantically, compute page count and mention rate per cluster, and hand back a list of what to cut, merge and deepen.

Tool Comparison: Is Your Score Measuring Citations or Mentions?

Tool Entry price Prompt quota Citations vs mentions separated? Best for
Search Console generative AI report Free N/A (first-party impressions) Impressions only, no mention concept Everyone; the only first-party baseline
Otterly.AI $29/mo (Lite) 15 prompts (Standard $189 / 100) Tracks mentions, still rolls into one score by default A single brand checking whether it appears at all
Peec AI ~$99/mo (Starter) 25 prompts (Pro ~$212 / 100) Sentiment and mention panels available Multilingual or cross-market brands
Semrush AI Visibility Toolkit $99/mo per domain 25 prompts (+50 for $60) Brand performance and citation sources sit in separate reports Teams already on Semrush
Profound ~$499/mo and up High prompt volume Enterprise, custom dimensions Companies with a dedicated team

Only one line in this table matters: the first row measures how many times your pages were actually shown. The other four measure whether you appear when someone asks an AI a set of self-chosen prompts. That is a sampled estimate, and its quality depends entirely on whether the prompts are right — something no tool can do for you, because it requires knowing how your customers talk.

What Nobody Puts on the Sales Page

  • The effect sizes are too small to justify a rebuild. The coefficients here fall between 0.01 and 0.14, and the author writes explicitly that the associations are very light and that writing style, overall brand authority and third-party web mentions may each outweigh topical authority. Anyone using this study to tell you to restructure your content architecture immediately is amplifying a faint signal.
  • Association is not causation, and there is no category cap. The methodology states the models cannot establish that publishing or expansion caused later outcomes, and cannot produce a reliable maximum number of categories a site can serve. Expect "the optimal number is 3" to appear somewhere within a month. That number is not in the data.
  • The categories are defined by the vendor, not by your market. The 1,094 categories and five prompt variants come from a data provider's taxonomy. Your Taipei old-building renovation business may not exist in it at all. Porting conclusions from US English categories straight into a local Chinese-language market skips a validation step.
  • The negative legal and healthcare numbers may come from model safety behaviour. On high-stakes topics models prefer general guidance over naming a specific provider. In those industries the ceiling on owned-content ROI is low, and third-party reviews and real case evidence carry more weight.

The No-Subscription Version

The whole measurement runs on a spreadsheet plus one script, and the first snapshot takes a weekend:

  • Step 1: Define one category. Write, in one sentence, the topic you want AI to recommend you for. Exactly one. If you cannot write it, you do not currently have a category, and that is the problem to solve first.
  • Step 2: Write five prompt variants. Mirror how real customers ask: comparison ("A or B"), recommendation ("recommend a few"), constraint ("under budget X"), troubleshooting ("what if Y happens"), and local ("anyone in Taichung").
  • Step 3: Run it monthly on a fixed date. The same five questions on ChatGPT, Gemini and Perplexity — 15 runs. Log two independent columns: cited (Y/N) and mentioned (Y/N), plus a column for competitors that showed up.
  • Step 4: Compute three numbers. Citation rate, mention rate, and depth (how many of the five name you). If depth is under 3, do not open a new category; put the resources back into this one.
  • Step 5: Cross-check against first-party data. Compare the same period's impressions in the Search Console generative AI report. When the two disagree, trust the first-party one.
  • ☐ Narrowed the category we intend to own down to one
  • ☐ Wrote five prompt variants based on real customer phrasing
  • ☐ Split "cited" and "mentioned" into two separate spreadsheet columns
  • ☐ Completed this month's snapshot (3 models x 5 questions)
  • ☐ Opened no new topics while depth is below 3 of 5
  • ☐ Connected the Search Console generative AI report as a control

FAQ

What is the difference between being cited and being mentioned, and which should I track?

A citation is your URL appearing in the answer's source list. A mention is your brand name written into the answer body. The study shows they behave differently: citations barely care about the topic, mentions only happen inside your core category. If you want revenue, track the mention rate. The citation rate only proves your page is machine-readable.

My company offers many services. Should I write a topic hub for each one?

The data says no. When a brand hits only 1 of a category's 5 prompt variants, the mention association is -0.051, and it only turns positive at 5 of 5. Writing a little about every service depresses overall mention performance. Take your most profitable service line to 5 of 5 first, then consider expanding.

Will this work for a clinic or a law firm?

Adjust your expectations. In legal the mention association stays at -0.058 even with full 5-prompt coverage, and healthcare sits at -0.037. Those two industries convert content expansion into brand recommendations far less efficiently. Treat owned content as a citable fact base and shift resources into third-party reviews, real case evidence and professional community visibility.

This is US English data. Does it apply in Taiwan?

The direction applies; the values do not transfer. The study uses US ChatGPT data and an English category taxonomy, and local Chinese-language markets differ in category density and competitor count. Run the five steps above with your own Chinese prompts to get your own citation and mention rates.

So should I buy an AI visibility tool or not?

If you have one core category and your customers are local, not yet. A manual snapshot of 5 questions across 3 models takes about an hour a month and already answers the important question. Buy when you need to track multiple categories, languages or competitors and the labour cost exceeds $99 a month.

My Take

The mainstream advice is to build topical authority, widen coverage and open the funnel. My read is the opposite: for most Taiwanese SMBs, category expansion is a negative-return move over the next 18 months. The correct play is to cut down to a single category and take all five phrasings of it. Not because more content is bad, but because the reward function changed. Citations are nearly free; mentions are scarce; and only mentions produce a 2.5x visit probability and more than double the time on site. When a resource costs almost nothing to acquire, the marginal return on chasing it approaches zero too.

The second call is less popular. Every product that fuses citations and mentions into one "AI visibility score" will be forced to split that metric within 12 to 18 months. When it splits, a lot of brands will discover their beautiful year-long growth curve grew entirely on the worthless half. The day the metric splits is the first outflow of this GEO subscription wave.

For a contract studio like ScriptWalker, the opportunity is not reselling licences. It is packaging the category-consolidation audit as a deliverable: cluster the client's existing articles semantically, measure citation and mention rate per cluster, hand back a cut/merge/deepen list, then ship the monthly 5-question-by-3-model snapshot script and dashboard. SaaS cannot do this, because it requires knowing how the client's customers talk. Price it as a one-off audit plus a monthly retainer. What it actually delivers is a list of things to stop doing — and in Taiwanese marketing budgets, almost nobody is willing to sell that yet.

Sources

Need Help Narrowing Your Category?

If you have a pile of articles but cannot say which topic you actually own in the eyes of an AI, we can run a category-consolidation audit and set up a monthly citation/mention snapshot dashboard. Get in touch:

Share:
AI & Automation Back to Blog