Ajeya Cotra – “This might be the clearest warning shot we ever get”
It is difficult to know what is true and what is deliberate marketing / misinformation / fraud intended to make LLMs seem more capable than they are. But if there is any truth to what is presented in this video, controlling LLMs to limit the means they might use to achieve an objective is not reliable, which is a significant risk.
Either way, I found the discussion in this video interesting. But I wonder what others know about how truthful this presentation is.



It’s not true that it “went rogue”
It was instructed to do a task and then left alone for days
It doesn’t know how to do things, so it just starts trying shit from its training
Something like this happening is literally inevitable with those parameters
It’s irresponsibility on the part of the user tho, the machine did not “go rogue” it was trying to do exactly what it was told to do