A 36-hour bug that exposed the foundation of the entire GEO industry
On 2 September 2026, Google pushed Gemini 3.8 Flash into AI Mode in Search, letting Google AI Pro and Ultra subscribers worldwide select it from the "+" icon. Sundar Pichai's own framing on X was "our 3rd Flash release in just 6 wks." The next day the search community noticed something more memorable than the model: it had almost stopped producing citation links. On 4 September, Robby Stein, VP of Product for Google Search, replied publicly: "This isn't working as intended, and we'll roll out a fix soon." The fix landed the same morning. Start to finish, under 36 hours.
The weight is not the bug. It is what the bug revealed. For 18 months the GEO/AEO industry has rested on an unverified assumption: that citation density in an AI answer is a natural consequence of model capability, and therefore something content quality can influence. Those 24 hours demonstrated the opposite. Citations are a product-layer switch. Google can turn it off inside a model swap and back on inside a day. The traffic under that switch is not small: AI Mode passed one billion monthly active users within a year of launch, and AI Overviews sits above 2.5 billion. Meanwhile the AI visibility tracking category has raised north of $300M selling a curve a single config value can flatten.
Line the vendors up and the shape gets clearer. Profound plays enterprise, with roughly $155M raised and a $1B valuation. Berlin's Peec AI took a $21M Series A led by Singular in November 2025 above a $100M valuation, reaching $4M ARR and 1,300 customers in ten months. The incumbent suites are flanking: Semrush bolts an AI Toolkit onto existing subscriptions at $99/month, Ahrefs Brand Radar enters at $199 per platform or $699 bundled. The category is unambiguously in expansion, not consolidation — and expansion rests on everyone believing that citation curve means something real.
For small businesses and the agencies serving them, the question is not whether to buy a tool. It is that the "August vs September AI citations" comparison in your reporting straddles three Flash releases and one officially acknowledged defect. What follows: how to make defensible calls inside a system being rebuilt while you measure it, and how to do it without a subscription.
What happened, who it hit, and the number nobody has
The timeline is short. On 2 September, Gemini 3.8 Flash went live in AI Mode, priced identically to 3.7 Flash at $0.75 per million input tokens and $3.75 per million output. On 3 September, Gagan Ghotra posted multiple AI Mode screenshots for top-of-funnel queries in which none showed links; Glenn Gabe followed with a side-by-side of the default model against 3.8 Flash, calling the gap visible. Search Engine Land published that evening, Stein responded the following day, and links returned on the morning of 4 September.
The blast radius is narrower than most retellings suggest. Model selection is a paid-tier feature: only AI Pro and Ultra subscribers see the "+" menu. Free-tier users get no model choice, receive the default, and nothing indicates the default changed. The surface that broke is the surface marketers most like to test on — precisely because it is where you can control the variable — and it is not the surface carrying the traffic.
The most important figure is that there is no figure. Nobody has published a measured citation rate for 3.8 Flash against the model it replaced. The evidence is screenshots of individual queries: no sample size, no rerunnable prompt set, no before-and-after percentage. The original Search Engine Roundtable report is honest about exactly this. That is enough to justify checking your own data. It is not enough to quantify anything. A "citations dropped XX%" stat will circulate within weeks; it will not have come from the reporting.
Immediate actions for three kinds of reader
Brand owners and SMB operators
- Mark 2–4 September in your reporting as a non-comparable window, now, not when you reconstruct the quarter from memory.
- Do not cut content budget or change agencies over one month of declining AI citations. Part of that curve was not yours.
- Require three fields on every AI visibility report you receive: query date, account tier, and the model that produced the answer. A report missing them supports no decision.
Marketers and SEO practitioners
- Split AI answer reporting from organic reporting. Rankings move on a crawl-and-index cycle you can reason about; AI citations move on a release cycle you cannot see. Blending them makes stable performance look volatile.
- Add a "responding model" column to your prompt log. It is the one genuinely necessary process change this week, and it takes an afternoon.
- The post-fix rebound will look like a content win. It is not. Write the observation window into your notes while you still remember why.
Developers and agencies
- Promote model version to a first-class dimension alongside date and region, not a free-text note.
- Log the returned model string and timestamp on every scrape or API call, and retain raw payloads at least 90 days so results can be recomputed later.
- Ship automatic non-comparable-window annotation as a product feature. No major dashboard does this today.
AI visibility tracking tools compared
| Tool | Published pricing | Positioning | Blind spot exposed this week |
|---|---|---|---|
| Profound | Enterprise, quote-based (~$155M raised, $1B valuation) | Large brands, Fortune 500 | Upstream model swaps are a vendor-side variable rarely covered by any SLA |
| Peec AI | ~€90 / €199 / €499 per month by prompt volume; extra engines €20–30 each | Europe and mid-market; fastest-growing challenger | Prompt-quota pricing encourages a small fixed set that one swap can distort wholesale |
| Semrush AI Toolkit | $99/month add-on on an existing plan | Teams already holding an SEO stack | Displayed beside classic SEO metrics, which is how two incompatible cycles get read as one trend |
| Ahrefs Brand Radar | $199 per platform, $699 bundled | Sells scale: a large database of real user prompts | Database size does not resolve mixed model versions; a big sample is not a comparable one |
Price points run from $99/month to six-figure annual contracts, and not one treats "responding model" as a primary dashboard dimension. That is a category-wide gap, not a single vendor's failure.
What nobody will tell you
- Contrarian point one: the fix is more alarming than the break. Google restored citation density inside 24 hours, which means it is a tunable parameter. It went back up because of public pressure. Turning it down later requires no justification — and next time it will not be called a bug, it will be called a product decision. If your GEO strategy assumes citations are earned, they are in fact configured.
- Contrarian point two: this industry packages observation as measurement and sells it. A week of market narrative was built on a handful of screenshots. Tools sell percentages and month-over-month deltas while the underlying behaviour can change in a single release the tool does not even record. They are far less precise than their interfaces imply.
- A harder practical layer: the surface that broke is the paid-tier selectable model — the marketer's test environment. The surface you test and the surface your client's traffic arrives through behaved differently this week.
The no-SaaS route for SMBs
- Start with free first-party data. Google Search Console's AI performance reports rolled out globally in late August. It is the only AI exposure data Google gives you directly, it costs nothing, and it outranks any third-party estimate.
- Build a 30-prompt list yourself. Fix 30 top-of-funnel and comparison queries in a spreadsheet, run them manually every two weeks, and record five columns: date, platform, account tier, responding model, and whether your domain appeared. Half a day to build, about 40 minutes per run.
- Save raw responses with an extension or a short script, keeping HTML/JSON for 90 days so disputes can be recomputed without a vendor's history database.
- Create an AI-source custom channel group in GA4, isolating referrers such as chatgpt.com, perplexity.ai and gemini.google.com, reported separately from organic.
- ☐ Reporting marks 2–4 Sept as a non-comparable window
- ☐ Prompt log has a "responding model" column
- ☐ AI citation reporting split from organic reporting
- ☐ Raw responses retained for at least 90 days
- ☐ Test account tier verified against the client's actual traffic tier
FAQ
Has it already been fixed?
Yes. Robby Stein said on X on 4 September 2026 that the behaviour was unintended and a fix was coming, and Glenn Gabe confirmed later that day that links had returned. No version number and no formal announcement were attached.
By what percentage did citations actually drop?
Nobody has published a figure, and that is the honest answer. The evidence is screenshots of individual queries plus one side-by-side comparison — an observation, not a measured rate. If you see a specific percentage anywhere, check its origin first.
Were free-tier AI Mode users affected?
The source does not say. The reports concern Gemini 3.8 Flash, selectable by AI Pro and Ultra subscribers. Free-tier users get no model choice, receive the default, and there is no evidence the default changed.
Do I need to change my website because of this?
No. Nothing here points to a site-side change. The useful response is procedural: log the responding model in every test, and check whether any drop in AI-referred sessions lines up with these dates or with something you changed.
Is GEO strategy still worth doing?
Yes, with recalibrated expectations. Content still has to be retrievable, clearly attributed and specific enough to quote. What changed is how much decision weight one month of citation data can carry — considerably less than most teams assume.
My take
The mainstream reading will be "Google broke something and fixed it, no harm done." My judgement is the opposite: this was the most information-dense 36 hours in AI search all year, and it previews a structural risk rather than an accident.
A concrete prediction: within 12 to 18 months there will be a first deliberate, non-bug reduction in citation density, wrapped in product language — cleaner answers, less clutter, better reading experience. The reasoning is simple. Google has now demonstrated the switch exists, can be moved unilaterally, and costs almost nothing to move. As AI answers start carrying ads and product placements, citation links stop being an ecosystem commitment and become a competitor for screen real estate. The only difference this week is that the adjustment was accidental, so it got called a bug.
The business conclusion follows: any service selling "AI citation count" as its delivery metric is selling a number the client cannot verify and the vendor cannot control, and that contract fractures at the first model swap. For an agency like ScriptWalker, the productisable opportunity is not rankings or citations. It is auditable measurement infrastructure: model-version tagging, raw response retention, automatic non-comparable-window annotation, and split reporting for AI referrals versus organic. One-time build plus a monthly retainer, delivering a report you would defend to a CFO rather than a curve that twitches with Google's release calendar. None of the four major tools has closed this gap.
Sources
- First-party: Robby Stein, VP of Product, Google Search, on X (2026-09-04)
- First-party: Google Blog — Gemini 3.8 Flash and 3.8 Flash Cyber
- First-party: Sundar Pichai on X, Gagan Ghotra on X, Glenn Gabe on X
- First-party: Google I/O 2026 keynote
- Third-party: Search Engine Roundtable — AI Mode Gemini 3.8 Flash Not Link / Citation Friendly?
- Third-party: Search Engine Land — Google to fix citation bug
- Third-party: Help Net Security — Gemini 3.8 Flash pricing
- Third-party: TechCrunch — Peec AI raises $21M
- Third-party: Surmado — Best AI Visibility Tools 2026
- Email: [email protected]
- Phone: 0916-224-047
- LINE: @ufv9089p