Forget being remembered. In June 2026, Jeff Oxford of Visibility Labs ran 1,000 product-recommendation prompts through ChatGPT — 20,000 responses in total, half with search switched on, half with it off. Only 19.8% of the brands ChatGPT recommended from memory still made the cut once it could check the live web. Whatever goodwill your business built in its training data is worth almost nothing the moment a real customer asks a real question. Here's that study, a second one that makes it worse, and what actually survives the flip.

The 80.2% number

Oxford's method was deliberately boring, which is why the result lands: pull 1,000 real transactional keywords, turn each into "what is the best [X]?", run it 10 times with ChatGPT search enabled and 10 times with it disabled, and count what overlaps.

Product recommendations: search off vs. search on

SOURCE: JEFF OXFORD, VISIBILITY LABS — 20,000 RESPONSES, JUN 2026

SURVIVED THE TOGGLE

19.8%

What ChatGPT "knew" from training and what it says once it can actually look.

CHANGED COMPLETELY

80.2%

Even the brands ChatGPT recommended in every single no-search run only survived 15.8% of the time with search on.

Search-enabled answers were also shorter — 5.2 products on average versus 6.2 without search. The live web doesn't just reshuffle the list. It shortens it, which means fewer slots, contested harder.

If your entire AI-visibility plan was "get mentioned enough that the model remembers me," that plan just lost four-fifths of its value. Memory is the opening bid. The live web makes the final call.

Being in the training data gets you into the room. It doesn't get you the recommendation — the search does that, fresh, every single time someone actually asks.

The part that's actually good news

Read the same number the other direction: if 80% of the deck reshuffles on every real query, nobody's position is locked in — including the competitor you assume already owns your category. An incumbent's spot isn't inherited from last quarter. It's re-earned, or lost, on every single search. Instability is bad for whoever's currently winning and comfortable. It's the best news a currently-invisible business has had all year.

There's one number in the study that says what actually survives the churn: a 0.4 Pearson correlation between how often a brand shows up in ChatGPT's cited sources and how often it gets recommended. Not the strongest correlation you'll ever see — but in a dataset where 80% of everything else is noise, being a source the model keeps citing is the signal that doesn't wash out.

It's not even the same answer for two different people

Three weeks after Oxford's study, a second one made the picture worse. "The Personalization Gap," published September 23, 2026 by Cassie Wilson Clark and Joao da Silva, tested ChatGPT, Gemini, Claude and Perplexity across six user personas, three product categories and thirty prompts — comparing a clean session against one carrying real user history.

Brand-set divergence caused by user history alone

SOURCE: WILSON CLARK & DA SILVA, "THE PERSONALIZATION GAP" — 8,609 RESPONSES, SEP 23 2026

33.3ppgeminiLARGEST EFFECT MEASURED
16.0ppchatgptABOVE NORMAL VARIATION
7.2ppperplexitySMALLER, STILL PRESENT
0claudeNO DETECTABLE EFFECT

Same question, same brand, same day — different answer depending on what the model already knows about who's asking. There is no longer a single "your ranking in ChatGPT" to check once and file away.

We wrote three weeks ago that AI doesn't rank you, it remembers you — that repetition beats ranking. Still true. What these two studies add is the uncomfortable second half: memory gets you considered, not chosen, and even "chosen" isn't one fixed answer anymore. It's a live recalculation, personalized, every time.

What that costs you if you wait

None of this is abstract. Adobe Analytics, tracking more than a trillion visits across 200-plus top US retailers, found AI-referred traffic converting 54% better than every other channel combined by May 2026, growing 138% year over year. That traffic exists right now. It is being routed to whoever the model decides to name, on a system that reshuffles 80% of its answers per query and shifts by up to 33 points depending on who's asking. Every day you're not part of that recalculation, it isn't skipping the slot — it's filling it with someone else's name.

⚠ WHAT WON'T FIX THIS

A one-time "AI optimization" push. There is no state to reach and hold — the deck reshuffles per query, not per quarter. Chasing a single "ChatGPT ranking." There isn't one; there are four platforms, personalized, diverging by up to 33 points from each other. Waiting for the market to settle. It isn't going to. Volatility is the permanent condition, not a phase to wait out.

What actually holds up

Three things, in order of how directly the data supports them:

  1. Be a source worth citing, not a brand hoping to be recalled.

    The 0.4 correlation between citation frequency and recommendation frequency is the one number in this piece that isn't noise. Publish the specifics — prices, service areas, answered questions — that a model can lift and attribute.

  2. Check more than once, and check more than one surface.

    A single "we're doing fine in ChatGPT" check from three months ago tells you nothing about today, and nothing about Gemini, where the personalization effect is twice as strong.

  3. Stop optimizing for a fixed answer that doesn't exist.

    Optimize for being a stable, citable, quotable source instead — the one variable in this whole mess that the data says holds its value when the toggle flips.

Whatever your number is today, it isn't going to be your number next month, on this platform or on Gemini, for this customer or for the next one. The only response that survives that fact is to stop treating AI visibility as a box to check once.

The one-sentence version

  • 80.2% of what ChatGPT recommends changes the instant it's allowed to search — memory alone isn't enough
  • Even "always recommended" brands survived the toggle only 15.8% of the time
  • Recommendations are personalized by user history too — up to 33.3 points of divergence on Gemini
  • The one signal that correlates with getting recommended anyway: being cited, consistently, as a source
  • Instability cuts both ways — nobody's spot is locked in, including whoever currently has yours

Check where you actually stand before you plan around a number that's already stale: run the free scan, or read the 10-prompt check and run it across all four platforms — not just once.