We are exposed to a much narrower range of knowledge through AI chatbots than through a basic web search. This is documented by a new study led by the University of Copenhagen. As more and more of us turn to AI chatbots for information, researchers warn that the risk of ‘knowledge collapse’ increases.
Why so?
The exact same mechanism is on display here. Ask an LLM for information, and it’s more likely to condense an answer, not display additional relevant information, and slowly but surely produce less and less outputs including that given information.
Sources then begin to be generated by the AI model that adds less and less of that information over time, LLMs cite those new sources that don’t include that information, etc, etc, and now that information is just not present in a majority of sources, and is never stated in any LLM responses.
If anything, this is more likely than models forgetting certain vocabulary, because while the underlying tokens for a lot of that vocabulary still exist in the model, the subsequent combinations for a given piece of information are substantially less likely, and the models’ finite size can only store a subset of that actual information in statistical averages. (i.e. you can store the alphabet and a dictionary of every word in the English language, and many simple phrases, in a tiny fraction of what the sum of Wikipedia is, and Wikipedia itself will contain a subset of all information on the internet, and so on, but you can still theoretically write any English content you’ll find on Wikipedia or the internet using that alphabet and dictionary)