• AmbitiousProcess (they/them)@piefed.social
    link
    fedilink
    English
    arrow-up
    1
    ·
    8 hours ago

    I would be concerned about language collapse (i.e. models incrementally forget vocabulary + fail to adapt to new terms leading to language becoming incrementally worse) before I would worry about knowledge collapse

    Why so?

    The exact same mechanism is on display here. Ask an LLM for information, and it’s more likely to condense an answer, not display additional relevant information, and slowly but surely produce less and less outputs including that given information.

    Sources then begin to be generated by the AI model that adds less and less of that information over time, LLMs cite those new sources that don’t include that information, etc, etc, and now that information is just not present in a majority of sources, and is never stated in any LLM responses.

    If anything, this is more likely than models forgetting certain vocabulary, because while the underlying tokens for a lot of that vocabulary still exist in the model, the subsequent combinations for a given piece of information are substantially less likely, and the models’ finite size can only store a subset of that actual information in statistical averages. (i.e. you can store the alphabet and a dictionary of every word in the English language, and many simple phrases, in a tiny fraction of what the sum of Wikipedia is, and Wikipedia itself will contain a subset of all information on the internet, and so on, but you can still theoretically write any English content you’ll find on Wikipedia or the internet using that alphabet and dictionary)