On July 15, 2026, a critical survey uploaded to the Information Retrieval section of arXiv under identifier 2607.14035 quietly dismantled the one line the entire GEO (generative engine optimization) industry loves to repeat. Author Olivier Martinez systematically reviewed 45 studies published between November 2023 and July 2026, grading each by evidentiary weight. The conclusion is cold: no reviewed GEO technique produces a stable, cross-platform effect on whether a page is retrieved at all, or on actual downstream traffic. More striking is a counter-intuitive number: in an end-to-end experiment called SAGEO Arena (171,003 documents, 2,700 queries), rewriting only a page's body copy lowered its post-reranking top-10 presence by roughly 16%. You think you are optimizing; you are actually subtracting.
Why does this matter now? Over the past 18 months, GEO went from an academic term to a budget line item. AI chatbot referral traffic exploded from under a million monthly visits in early 2024 to more than 230 million by September 2025, while Adobe research found 98% of marketers lack a clear, confident AI-optimization roadmap. Tool vendors filled the vacuum, and nearly all of their pitches trace back to one figure: "GEO boosts visibility by 40%." That number comes from the 2024 SIGKDD foundational GEO paper — but the survey shows the 40% is really a single metric, Position-Adjusted Word Count, rising from 19.3 to 27.2 in a fixed setup where five documents were already fed to the generator. It has nothing to do with 40% more people clicking you.
Line up the peers and the context sharpens. AEO/GEO tracking tools like Profound, Evertune, Scrunch and Semrush have collectively raised over $200M in the past year, with 33 AI-citation trackers now on the market. What they all sell is the promise of being "citable by AI" — and the survey points precisely at the gap: experiments can show that a document already inside the context changes how it is cited, but rarely show whether that page gets retrieved at all, let alone whether it drives clicks or conversions. The category is not consolidating; it is scaling on an unproven premise.
Will small businesses and freelance studios be affected? Yes — and it is good news. If an expensive GEO subscription mostly buys you a lab metric, you probably do not need it. This article covers what the survey specifically overturned, which two levers actually work, what three reader types should do today, and how to do "get cited by AI" right without a single dollar of SaaS subscription.
Event details and the full numbers
This 18-page survey, with 8 tables, does not hand you tactics — it hands you a ruler for evidence. It decomposes "AI visibility" into a vector: discoverability, in-context exposure, citation probability, actual contribution, fidelity, and the final click or conversion. Its argument is that a technique can improve one component while damaging another, so collapsing them into a single score lets a promotional metric masquerade as a business result. Key findings:
- Rewriting body copy backfires: in SAGEO Arena, optimizing only the body cut top-20 presence about 9%, post-rerank top-10 presence about 16%, and final citation about 6%. The mechanism is simple: you raise the probability of citation given retrieval, but lower the probability of retrieval — net negative.
- Generic tips travel poorly: in C-SEO Bench (two tasks, six domains, ~1,900 queries, 16,360 documents), only 3 of 54 method-domain combinations were significantly positive, and none in question answering; in e-commerce, 10 of 15 heuristics were neutral or negative.
- Only two old things hold up: a factorial experiment of 252,000 trials across 6 models and 18 factors found query-document relevance and position within the context to be the primary determinants of first citation; other tricks have small marginal effects.
- Citation data is highly unstable: Google AI Mode swaps 56% of its cited sources weekly and ChatGPT swaps 74% — any single-week citation screenshot is a shaky basis for strategy.
- Traffic evidence is the weakest link: one controlled study found ChatGPT referrals grew 5.7x raw, but after controlling for platform growth the attributable effect was 1.82x (95% CI 1.31–2.54), with a conservative placebo test coming back inconclusive.
What three reader types should do now
- Brand owners / SMB founders: stop paying for subscriptions that "guarantee AI citations." Put the budget back into the two proven levers — content relevance to user queries, and your rank in traditional search. An independent analysis cited by the survey ranks "URL accessibility" and "search rank" as the two highest predictors of AI citation (9.5 and 9.4). You should be doing both anyway.
- Marketers / SEO practitioners: stop chasing a single-week citation score. Measure across engines, across query phrasings, with an untreated control group, and state whether you are measuring a rewrite effect once a document is injected, or real traffic. Add verifiable, dated, properly attributed facts (prices, definitions, statistics) — one of the few moderately effective moves the survey endorses, provided the numbers are real. Fabricated stats raise reuse but degrade answer accuracy.
- Developers / freelance studios: this is a productizing moment. Instead of selling "GEO magic," sell a machine-readable content foundation: clean information architecture, crawlable URLs, schema.org structured data, fast pages, clear source attribution. Bundle a "cross-engine visibility + untreated control" dashboard — more honest than selling unproven exposure, and harder to commoditize.
GEO / AEO tracking tool comparison
| Tool | Positioning | Rough price | Practical caveat |
|---|---|---|---|
| Profound | Enterprise AI brand-visibility tracking | Enterprise (high) | Pretty charts, but mostly "citation share" — mind the denominator trap |
| Evertune | Model-level brand-mention analysis | Enterprise | Cross-model comparison is useful; single-week data still unstable |
| Scrunch | AI search visibility monitoring | Mid-high monthly | Good for ongoing monitoring; not "optimize and it rises" |
| Semrush (AI tracking module) | Legacy SEO plus AI-citation tracking | Add-on to existing plan | Integrates with existing SEO workflow; better value |
Shared blind spot: these tools are good at measuring how often you are cited, but cannot guarantee that optimizing makes it rise. The survey's warning is exactly this — do not treat a dashboard's citation share as proof that traffic will follow.
What they won't tell you
1. "Cited" does not mean "accurately represented." The fidelity studies the survey reviews found that in early generative engines only 51.5% of sentences were fully supported by their cited sources, and 74.5% of citations actually supported the claim they were attached to. Being cited may point to a page that does not support the statement — or to AI-generated content feeding back into the system.
2. Each engine cites a different source list. An audit of 4,706 queries found 53% of domains cited by Google AI Overviews were not in the organic top 10 and 27% were absent from the top 100; URL overlap between organic Google, AIO and Gemini was just 0.11–0.18. There is no single global ranking to optimize for — track one platform and you see about a third of your visibility.
3. White-hat optimization and manipulation share one channel — both alter the text an engine consumes. The survey documents that indirect injection can lift a target about 3 ranks on one commercial search API, and preference-manipulation attacks raised the recommendation rate of a fictitious camera from 34.0% to 59.4%. Aggressive tactics that work today may be flagged as spam tomorrow, with penalties compounding across traditional and AI search at once.
A no-SaaS-subscription alternative for SMBs
- Fix retrieval first, not body copy: make sure your pages are crawlable (robots, status codes, readable HTML) — the survey's top predictor, and completely free.
- Feed facts cleanly with structured data: use schema.org structured data to mark up article, author, date, price and definitions so engines can read verifiable, dated facts.
- Build your own cross-engine spot-check sheet: in a Google Sheet, log the same query set and manually ask ChatGPT, Gemini and Perplexity once a week, recording whether and how you are cited — free human intelligence that reproduces the core value of an expensive dashboard while exposing cross-engine gaps.
- Trust Google's own line: Google's official documentation states that optimizing for AI search is optimizing for the search experience — good SEO is the foundation of AI visibility, no separate magic required.
FAQ
So should I ignore GEO from now on?
No. GEO is a real phenomenon — it is true that already-retrieved content can influence an answer; it is unproven that rewriting boosts retrieval and traffic. Investing in relevance and traditional rank happens to nail GEO's two genuine levers.
What exactly is wrong with "GEO boosts visibility 40%"?
The 40% is a single metric (Position-Adjusted Word Count) rising in relative terms inside a lab setup where the document was already fed into context. It does not mean 40% more people click you, nor that you are more likely to be retrieved. The survey places it in its lowest confidence tier and rejects it as a general claim.
Do SMBs need tools like Profound or Scrunch?
Mostly no. They are good at measuring, not at guaranteeing growth. If you must track, pick one that integrates with your existing SEO workflow, or build a free cross-engine spot-check sheet for better ROI.
Does stuffing prices and statistics into pages help?
It has a moderate effect, but only if the numbers are real, dated and properly attributed. The survey warns that fabricated statistics raise reuse while degrading answer accuracy, hurting brand trust long term.
My take
The mainstream narrative says "AI search is here, hurry and buy a GEO tool to optimize." My call is the opposite: over the next 12–18 months most standalone GEO tools face a valuation correction, because their core promise — "optimize and you get cited, and traffic rises" — is being systematically contradicted by the academic evidence. When a category raises over $200M yet cannot produce stable causal evidence that it drives a single click, the smell of a bubble sets in. For a freelance studio like ScriptWalker, that is a clear productizing opening: don't sell GEO magic, sell a machine-readable, verifiable, measurable content foundation — clean technical SEO, structured data, fast pages, plus an honest cross-engine measurement dashboard. When the market finally accepts that the 40% was a lab number, the studios that deliver on substance rather than slogans will inherit the clients who got over-promised once already.
Sources
- arXiv 2607.14035 — Optimizing Visibility in Generative Engines: A Critical Survey of GEO (Olivier Martinez, 2026/07/15) (first-hand)
- arXiv HTML full text — full survey content and data tables (first-hand)
- arXiv 2311.09735 — GEO: Generative Engine Optimization (Aggarwal et al., foundational paper, source of the 40% figure) (first-hand)
- Google Search Central — AI features and your website (official stance: AI optimization is search optimization) (first-hand)
- PPC Land — Survey of 45 studies finds GEO rewrites can cut a page's AI retrieval 16% (2026/07/20)