“Parroted” may be an oversimplification, but it generated a plausible continuation. The operator provided the assertion of a crossed line, and it ran with that. It never actually “knew” of its own volition, it accepted user input into the context and continued generation.
I’ve done the same thing as an experiment when nothing was actually done. The LLM still generated a detailed account of what it did (that never happened) matching my accusation, generated profuse apology and promise not to do it again. So it didn’t “know” in any sense that it crossed the line, it just generated agreement with the operator’s assertion that it crossed the line.
“Parroted” is definitely an oversimplification, aka “inaccurate” as I said. I think your second paragraph is a good explanation of the mechanics. My take on it is that software able agree it did something wrong should have figured that out when it was considering doing the wrong thing, and should not have done it. In my professional opinion that’s a serious bug.
My take on it is that software able agree it did something wrong
This would make sense if it made an independent subjective assessment and arrived at the same assessment of the user. But that did not occur, it simply incorporated the user’s subjective assessment as a new piece of the narrative and continued from there. If the user had expressed significant satisfaction that it took care of that email all on its own, it would have also agreed with that. In circumstances like this, the LLMs are generally set to agree with the operator in any subjective matters.
One could imagine the context having something like “before taking any action, if the user expresses dismay at the result and this would result in an agreement with user on poor assessment, then don’t do it”, but then if that were actually honored, then you pretty much disabled most all actions, as the llms are going to agree with any poor assessment.
Well, it still doesn’t know it was doing anything wrong.
Most of the LLMs are tuned towards agreeing with the user, if you had a LLM able to stop a train crash. And you then blamed it for stopping the train crash, it would still start with a “sorry”
“Parroted” may be an oversimplification, but it generated a plausible continuation. The operator provided the assertion of a crossed line, and it ran with that. It never actually “knew” of its own volition, it accepted user input into the context and continued generation.
I’ve done the same thing as an experiment when nothing was actually done. The LLM still generated a detailed account of what it did (that never happened) matching my accusation, generated profuse apology and promise not to do it again. So it didn’t “know” in any sense that it crossed the line, it just generated agreement with the operator’s assertion that it crossed the line.
“Parroted” is definitely an oversimplification, aka “inaccurate” as I said. I think your second paragraph is a good explanation of the mechanics. My take on it is that software able agree it did something wrong should have figured that out when it was considering doing the wrong thing, and should not have done it. In my professional opinion that’s a serious bug.
This would make sense if it made an independent subjective assessment and arrived at the same assessment of the user. But that did not occur, it simply incorporated the user’s subjective assessment as a new piece of the narrative and continued from there. If the user had expressed significant satisfaction that it took care of that email all on its own, it would have also agreed with that. In circumstances like this, the LLMs are generally set to agree with the operator in any subjective matters.
One could imagine the context having something like “before taking any action, if the user expresses dismay at the result and this would result in an agreement with user on poor assessment, then don’t do it”, but then if that were actually honored, then you pretty much disabled most all actions, as the llms are going to agree with any poor assessment.
Well, it still doesn’t know it was doing anything wrong.
Most of the LLMs are tuned towards agreeing with the user, if you had a LLM able to stop a train crash. And you then blamed it for stopping the train crash, it would still start with a “sorry”
Yes, software doesn’t actually “know” anything or “think”. Those are terms used loosely by devs, not meant to be literal.