Reddit fell out of ChatGPT's citations, not out of the model
Klaas Foppen posted a chart on 18 August and called Reddit “almost wiped” from ChatGPT’s sources.
The data behind it has Reddit at a steady 3.8% of ChatGPT Search citations from 18 July to 7 August. On 14 August the share dropped under 1%, and the 14 to 17 August average is 0.5%.
That’s an 86% drop in a week, based on millions of real ChatGPT UI responses.
The replies split between celebration and panic.
I think both sides are misreading what that chart actually measures.
It measures citations.
That is not the same thing as what ChatGPT knows about Reddit, what it may have learned from Reddit, or whether Reddit can influence which brands the model recommends.
The retrieval system changed
Start with the timing.
When you ask ChatGPT certain questions, it can generate its own search queries before answering.
On 8 August, six days before Reddit’s citation drop, site: queries jumped from 0.37% of the query mix to 16.8% in a day.
A site: query is interesting because the model is effectively deciding which domain it wants to search before retrieval happens.
That changes the opportunity for domains to appear incidentally.
With an open web search, Reddit might rank alongside publishers, review sites and vendor pages.
With something like:
best running shoes site:reddit.com
Reddit was selected before the search even happened.
But if ChatGPT instead generates:
best running shoes site:runnersworld.com
Reddit never gets the chance to compete in that particular search.
So as more fanout queries become explicitly domain scoped, source selection increasingly happens before retrieval.
That matters.
I saw the search machinery changing in my own network captures too, although my measurement is different from Promptwatch’s fanout metric.
In my captures, the number of underlying search events I was counting fell from around 12 per answer to around 4.
Promptwatch, using its own fanout measurement across a much larger dataset, saw average fanout queries increase around the same period.
Those numbers aren’t directly comparable, but both datasets point to the same important thing:
ChatGPT’s search behaviour changed.
Some replies blamed Reddit itself because old.reddit.com started requiring logins in July, making anonymous access harder.
Possible.
But the timing of the fanout changes makes retrieval behaviour a much more interesting explanation.
Promptwatch’s own interpretation is that ChatGPT appears to have changed how it selects sources. They also haven’t completely ruled out a data collection issue on their side.
So I wouldn’t interpret this chart as:
ChatGPT suddenly decided Reddit is bad.
It looks much more like a retrieval change.
And retrieval settings change all the time.
Reddit’s ChatGPT citations collapsed before, in September 2025, then recovered.
OpenAI could change the fanout logic again next month and move the line in the opposite direction.
Citations are not the model
This is the part I think is getting lost.
OpenAI describes its foundation models as huge sets of numerical weights or parameters.
During training, the model processes enormous amounts of data and those parameters are adjusted as it learns patterns and relationships.
It doesn’t keep a searchable copy of every page it trained on.
The training changes the model.
That distinction matters here.
OpenAI says its foundation models are developed using publicly available internet data, information accessed through third party partnerships, and data provided or generated by users, trainers and researchers.
Reddit has been part of the public web for nearly two decades.
So publicly available Reddit discussions may already have influenced models through web training data, independently of the current Reddit partnership.
And once patterns and associations have been learned into model weights, the model does not need to retrieve the original Reddit thread to use them.
That is completely different from citations.
And then there’s the Reddit deal
OpenAI and Reddit announced their partnership in May 2024.
OpenAI gets access to Reddit’s Data API, providing real time, structured Reddit content that can help its products understand and surface Reddit discussions, particularly around recent topics.
But there’s an important distinction here.
The public announcement does not say:
We are taking the Reddit Data API and using everything in it to pretrain GPT.
OpenAI does say more generally that information obtained through third party partnerships can be used to train and improve its models.
But we don’t publicly know exactly how Reddit’s licensed API data is used across pretraining, post training, retrieval or other product systems.
So I’m not going to pretend we do.
We don’t need that assumption anyway.
The important point is simpler.
A drop in Reddit citations tells us something about retrieval. It tells us very little about what associations already exist inside the model.
And I’ve seen evidence of that distinction in my own experiments.
The brand can appear before the search
When I analysed ChatGPT’s network traffic for my shortlist research, I separated brands into two groups.
Brands ChatGPT wrote into its own first search query.
And brands that only appeared after web results were retrieved.
The difference was massive.
Brands already present in that first query were mentioned in 68.9% of the final answers.
Brands discovered only during retrieval were mentioned in 2.1%.
I also found 86 cases where ChatGPT recommended a brand without fetching that brand’s website anywhere in the conversation.
That is important.
Because it means retrieval isn’t necessarily where the shortlist starts.
Sometimes the model appears to begin with candidate brands already in mind and then searches around them.
Where those associations ultimately come from is much harder to prove.
My working hypothesis is that long term prominence across the open web helps determine which names the model reaches for before retrieval.
Reddit has historically contributed a huge amount of that discussion.
But I cannot look inside the weights and tell you:
This brand appeared because of these 427 Reddit comments.
Nobody doing AI SEO can.
What I can observe is what happens before and during retrieval.
And the distinction is pretty clear:
Being cited is not the same thing as being known.
Google barely moved
There’s another reason I wouldn’t suddenly write Reddit off.
Google didn’t react anything like ChatGPT.
The same dataset shows Reddit citations in Google’s AI surfaces changing much more gradually.
And outside AI citations, Reddit still ranks across traditional Google Search for loads of commercial queries.
Search for the kind of questions people ask before spending money:
best X
X vs Y
is X worth it
X reviews reddit
best X reddit
Reddit is everywhere.
Those discussions shape what people see before buying, regardless of whether ChatGPT puts a Reddit citation underneath its answer this week.
So is Reddit dead for AI SEO?
No.
But one Reddit tactic might be.
Pumping Reddit full of AI written posts purely to farm ChatGPT citations looks considerably less attractive if ChatGPT is retrieving and citing Reddit less often.
And Reddit was already attacking that problem itself.
The value of Reddit was never just the clickable citation underneath a ChatGPT answer.
It’s millions of people discussing products, comparing companies, complaining about software, recommending alternatives and arguing about what they actually buy.
Some of that content ranks in Google.
Some gets retrieved by AI systems.
Some may contribute to future model development.
And some of the historical public web may already have contributed to the associations models learned during training.
Those are four different mechanisms.
Lumping them all together as “Reddit citations” is the mistake.
So when a chart says Reddit citations fell 86%, the conclusion isn’t:
Reddit lost 86% of its value to ChatGPT.
The conclusion is:
ChatGPT cited Reddit 86% less during that period.
That’s what the data actually measures.
Everything else needs another experiment.
I’ll keep watching the fanout mix in my own captures, and anything that moves goes into the ChatGPT research tracker.
So for now, please,