AI & Automation

AI's Content Graveyard: 57% of Cited Domains Vanish After One Month, and the 2.7% That Survive Did One Thing Differently

2026.08.18 · 25 views
AI's Content Graveyard: 57% of Cited Domains Vanish After One Month, and the 2.7% That Survive Did One Thing Differently

Somantra's Aug 17 study of 2,437,107 ChatGPT and Google citations: survival is decided by content format, not domain authority — and pages labelled "complete guide" are 3.5x more likely to end up in the graveyard

Share:

Cited by AI once, and then nothing

On 17 August 2026, Sydney-based AEO platform Somantra published a study through GlobeNewswire that tracked 2,437,107 citation records across 28,725 domains on ChatGPT and Google over seven months. The finding: 57.2% of domains appeared in exactly one month and never surfaced again, while only 2.7% held a citation in all seven. Somantra calls the gap between those figures the content graveyard.

That this is only being measured now is no accident. For 18 months the AEO/GEO category has treated "how often does AI mention us" as the only scoreboard, and the commercial model was built on snapshots: fire 100 prompts, compute a mention rate, ship a dashboard. Capital followed. AI visibility tools raised more than $300M between summer 2025 and spring 2026; Profound closed a $96M Series C at a $1B valuation in February; Adobe completed its roughly $1.9B acquisition of Semrush on 28 April. Nobody sold an answer to "how long does the citation last." Somantra is the first to stretch the timeline until the snapshot becomes a survival curve, and that curve is brutally steep.

The peer set sharpens the gap. Profound, with over $150M raised, sells enterprise prompt monitoring. Scrunch went elsewhere on 13 August, showing AI referrals account for just 1.1% of news publisher visits. AirOps dug into retrieval and found ChatGPT fetched 548,534 pages across 15,000 prompts but cited only 15%. Three companies answering "who gets mentioned," "does the citation send traffic," and "why wasn't I cited." Somantra adds the fourth: once cited, does it hold.

For small and mid-sized businesses in Taiwan this is good news. Somantra states plainly that domain authority, brand recognition and publishing volume explained very little about which domains persisted; content format explained a great deal. Format is something a 15-person company can change this quarter. What follows: which formats survive, which one dies, and how to measure your own survival rate without a subscription.

The full numbers: what lived, what died

The study covers the Australian insurance category over a seven-month window, spanning citations on both ChatGPT and Google. Somantra flags it as correlation across a single category, not established causation.

  • 57.2% of domains were cited in a single month, then dropped to zero.
  • 2.7% held citations across all seven months.
  • Pages using discount or savings language appeared among long-term survivors at roughly twice the rate of one-citation URLs.
  • Comparison-format content showed a similar advantage; FAQ and how-to correlated positively with persistence.
  • Content labelled a "complete guide" was 3.5x more common among domains that vanished after one citation.

The surviving 2.7% is dominated by comparison sites, government explainers, and insurers whose pages lead with those formats. The losers are content teams whose budget sits entirely in long-form guides — precisely what won on Google in 2020. Founder Arun Prasad put it bluntly: those pages are cited once and then left behind.

The study also hands over a metric you can drop into a report: content survival rate, the share of pages still cited after three or more months, replacing publishing volume and traffic as a standing KPI. You do not need to buy their product to calculate it.

What each type of reader should do today

Brand owners and SMB founders:

  • Replace "how many pieces did we publish" with "of the pages published three months ago, how many are still cited."
  • Check whether your site has one page that is purely an "us vs competitors" comparison table. Most Taiwanese SMBs do not — the cheapest gap on the board.
  • Publish price bands, plan differences and common add-ons wherever you legally can. Pricing language was one of the strongest survival signals in the data.

Marketing and SEO practitioners:

Developers and agencies:

  • Add a scheduled Laravel command calling the ChatGPT and Perplexity APIs monthly with a fixed prompt set, writing cited URLs to the database.
  • Parse server logs to separate GPTBot, ClaudeBot, PerplexityBot and Google-Extended crawl frequency — OpenAI's official bot documentation lists the exact user agents. A page never re-fetched cannot be re-cited.
  • Redesign CMS content types so comparison tables, FAQ blocks and plan pricing are reusable components, not hand-written HTML.

AI visibility tool comparison

ToolPositioningEntry priceBest forWhat's missing
ProfoundEnterprise prompt monitoring plus agentic execution~US$295–399/moMid-to-large brands with a dedicated SEO teamToo expensive for Taiwanese SMBs; snapshot-first
Semrush AI ToolkitAI module inside an existing SEO suite~US$99/moTeams already paying for SemrushFree tier discontinued; costs stack fast
Peec AIMid-market mention tracking, ecommerce-leaningFrom ~€89/moEcommerce and D2C small teamsThin Chinese-language sampling
Otterly AILightweight entry-level citation tracking~US$29/moSingle-site, tight budgetsLow prompt quota; weak benchmarking
SomantraMindshare / Consideration / Engagement scoringEnterprise quote, free auditFinance, insurance, high-consideration categoriesUnknown Taiwan coverage

Shared blind spot: all five sell "how many mentions right now." None ships citation survival rate as a default report column — including Somantra, which published the study.

What nobody will tell you

Counterpoint one: format is a proxy, not a cause. Comparison tables, FAQs and pricing pages survive not because of what they are called but because they natively resolve one question in one extractable passage. Renaming a long article "X vs Y" without breaking it into independently extractable units will not reproduce the effect.

Counterpoint two: much of that 57% is not a content problem, it is re-crawl economics. Re-citation requires re-fetching, and AirOps found ChatGPT cited only 15% of what it retrieved — retrieval is already the narrowest part of the funnel. For SMB sites that update rarely, link internally poorly and never touch sitemap lastmod, the disappearance happens at the crawler layer rather than the ranking layer, and those two problems have completely different fixes.

One more limitation worth stating plainly: the sample is Australian insurance, one category, one seven-month window. Insurance is heavily comparison-shopped and dominated by authoritative government explainers, so "comparison tables and pricing language win" is close to a foregone conclusion there. Port it to Taiwanese restaurants, machine tools or B2B SaaS and the ratios will move.

The no-subscription approach

  • ☐ Create a Google Sheet: URL, publish date, format tag (comparison / FAQ / guide / pricing / case study), and twelve monthly cited-yes-no columns.
  • ☐ Fix 20–30 prompts your actual customers would type. Run them monthly in ChatGPT, Gemini and Perplexity. Twenty minutes a month.
  • ☐ Advanced: a Laravel command or Python script running the same prompts through official APIs into SQLite, fired by cron monthly.
  • ☐ Pull GPTBot, ClaudeBot and PerplexityBot hits from your access log and count monthly fetches per URL with awk. Rewrite whatever dropped to zero first.
  • ☐ Maintain lastmod per Google's official sitemap guidance so genuinely updated pages can be re-fetched.
  • ☐ Each quarter, split five atomic pages out of each of your three highest-traffic guides; compare survival rates three months later.

On the survival-rate column specifically, this equals a $300/month dashboard — because those dashboards do not have that column.

FAQ

The study only covers Australian insurance. Does it apply to Taiwanese restaurants, manufacturing or ecommerce?

The ratios do not transfer; the method does. Insurance is heavily comparison-shopped and regulated, so comparison tables and pricing language are structurally advantaged and the numbers get amplified. Content survival rate itself is industry-agnostic. Measure the same columns for three months and derive your own ratio.

Should we really stop writing complete guides?

Write them, but not as citation sources. Long-form guides still convert and work well as pillar pages. The atomic pages split out of them are what gets cited. The real mistake is putting 80% of your content budget into a ten-thousand-word article with no extractable passages.

How do I know if my pages are already in the graveyard?

Two cheap checks. If GPTBot and ClaudeBot have not re-fetched a URL in three consecutive months, it will almost certainly not be cited again. If your fixed prompt set has not surfaced your domain in three consecutive months, the page goes on the rewrite list. Neither needs a paid tool.

Is a $300 per month AI visibility tool worth it?

Yes if you have a dedicated marketing team, publish more than 20 pieces a month, and need competitor benchmarking for a board deck. No if you are under ten people publishing two to four updates a month — the information density is lower than a spreadsheet you maintain yourself, and the tool's history loses comparability every time a platform swaps models.

Taiwanese B2B firms cannot publish prices. What replaces the pricing-page advice?

The value of pricing language is not the number, it is the boundary it lets a reader judge against. Starting ranges, a three-plan difference table and what triggers a surcharge all give an engine an extractable, comparable passage.

My take

The mainstream reading will be "stop writing long articles, write comparison tables and FAQs." I think that conclusion expires within 12 to 18 months.

Format advantage is really retrieval-architecture advantage. Today's engines fetch live, extract passages and assemble an answer, so extractability dominates. But every major platform is moving toward persistent brand memory and entity indices. Once an engine holds a stable entity understanding of a brand, it stops reassembling answers from freshly scraped passages. When that happens, sites that climbed by renaming headlines and bolting on comparison tables will fall back together, and what remains standing will have a machine-readable entity footprint: complete structured data, consistent naming across platforms, and first-party data worth citing.

So my call: comparison tables and FAQs are short-term arbitrage worth doing, but they are not a strategy. The durable asset is whether your company reads as a clearly defined entity to a machine.

For an agency like ScriptWalker the productisation opening is concrete: package "AI citation survival audit" as a fixed six-week engagement — week one builds the prompt set and log parser, weeks two to three restructure content types, weeks four to six deliver a three-month tracking dashboard. The stack is a Laravel scheduler plus a light front end, and the client receives a column they cannot buy. Better repeat revenue than one-off rebuilds, better margin than reselling foreign SaaS.

Sources

Primary:

Third-party:

Share:
AI & Automation Back to Blog