Advanced AI models suffer a near-total collapse on classic psychology test as cognitive demands increase

sanitation@lemmy.today · 1 day ago

Advanced AI models suffer a near-total collapse on classic psychology test as cognitive demands increase

khornechips@sh.itjust.works · 19 hours ago

So… last week then?

Communist@lemmy.frozeninferno.xyz · 16 hours ago

I get that you hate AI but there’s no reason to lie about its capabilities.

Kay Ohtie@pawb.social · 10 hours ago

All of these features are not something the models themselves can do, but are grafted on.

I could easily write a Home Assistant automation pattern matching for nearly every way someone could say “how many Rs are in strawberry”, depluralize a plural letter, and run it against “wc” in a bash terminal.

That doesn’t mean it’s smarter. It’s that I’ve added something specific to it.

MCP and the like is just that too, gluing on functions or the ability to hopefully invoke a function. That’s why so many hilariously mundane ones exist.

At the core, it’s still a large language model: a statistical model of frequency of word and word chunk (token) patterns.

Sometimes one model can invoke another via that tooling but it’s still a grafting on. It isn’t a singular thing or system, but disjointed pieces so completely detached from how brains work.

This isn’t AI hate, it’s reality. I love the field of artificial intelligence and machine learning. It’s cool as hell. But an LLM is fundamentally incapable of being anything more than an LLM with glued on pieces that invoke functionality.

OpenAI saw people mock the inability to count so they wrote a specialized tool to count letters and glued it on.

The world is full of endless edge cases. The inability to simply resolve them without gluing on every single one means it just isn’t doing anything new.

MangoCats@feddit.it · 5 hours ago

I believe the progress of the last year is largely attributable to the appropriate “grafting on” of these wrappers around the LLM cores.

Communist@lemmy.frozeninferno.xyz · 9 hours ago

They regularly win olympiad mathematics up from not standing a chance and just created a novel solution to the erdos conjecture, them counting the r’s in strawberry is inconsequential but also something they can do even if you just use the raw api or a local model.

zbyte64@awful.systems · 5 hours ago

Using computers to search for a counter example to a conjecture isn’t exactly new ground and I suspect they did so with the aide of some harness tweaks like some numerical LSP. Like cool, it pushed the envelope but like what the parent said, they grafted on the ability to do a specific task.

criss_cross@lemmy.world · 13 hours ago

A lot of tools like Claude or ChatGPT have internal tools they call when they do math (or use a python script) rather than have the model actually compute anything.

The underlying tech itself can’t do it because you can’t do math by token probability.

Communist@lemmy.frozeninferno.xyz · 8 hours ago

Whether they use tools to do it or not is entirely unimportant, that’s just how they do it?

expr@programming.dev · 16 hours ago

That’s not lying. There’s nothing linguistic about numerical computation.

Communist@lemmy.frozeninferno.xyz · 8 hours ago

No.

https://www.nature.com/articles/d41586-025-02343-x

It’s lying

zbyte64@awful.systems · 5 hours ago

You know the “DeepMind and OpenAi models” is the hint that the LLM model is not the one doing the math. The LLM provides a hypothesis and the DeepMind model provides grounding or feedback on whether the hypothesis even makes sense or works.

Advanced AI models suffer a near-total collapse on classic psychology test as cognitive demands increase

Advanced AI models suffer a near-total collapse on classic psychology test as cognitive demands increase

Just a moment...