Ajeya Cotra – “This might be the clearest warning shot we ever get”

It is difficult to know what is true and what is deliberate marketing / misinformation / fraud intended to make LLMs seem more capable than they are. But if there is any truth to what is presented in this video, controlling LLMs to limit the means they might use to achieve an objective is not reliable, which is a significant risk.

Either way, I found the discussion in this video interesting. But I wonder what others know about how truthful this presentation is.

  • Fishnoodle@lemmy.world
    link
    fedilink
    arrow-up
    5
    ·
    5 天前

    Exactly. You have 12 months or less to cash out, and that’s being generous. It’s probably closer to 2-3 months

    Take your profit, leave a few shares if you want to hedge.

    You can’t daisy chain forever, and markets are going to find that out soon.

    • HakFoo@lemmy.sdf.org
      link
      fedilink
      arrow-up
      7
      ·
      5 天前

      At this point I can see parallels with Intel in the 2000s.

      If you remember then, they went all in on the Pentium 4/“Netburst” architecture. It was sort of a dog from day 1. It ran hot, it wasn’t very powerful, you needed purpose built power supplies and weird RAM at first. But it was the future because they planned to scale it to 10GHz which would make up for all its limitations. Sounds a lot like AI-- it’s expensive and still dodgy, but the planned future design solves it all.

      Narrator Voice: they never made it to 10GHz. Even at 3.8 it still sucked, and the thermals and power needs made it unviable to go further.

      The LLM industry is flogging their metaphoric 3.6GHz Pentium 4 now. Maybe you can get one or two more sales cycles out of the current paradigm, but not much more. If a costlier trillion parameter model is only marginally better than a 200B model (that was only marginally better than a 25B one…), how do the economics justify the hundred-trillion parameter model? Training costs likely rise nonlinearly, too. It becomes harder to find fresh content for the model. If we need to scan vintage out of print books, we must have exhausted the easier sources.

      Whether you’re a true believer really chasing AGI or just surfing the grift, you’re going to have to pivot hard soon if you have any business strategy that avoids commoditization. We’re reaching the largest economically feasible versions of the current design. Intel survived by having such a pivot in their pocket; the replacement Core series CPUs were originally the “B-team” designs for laptop chips quickly scaled for a new use case. Does any big-dollar AI platform have something that paradigm-breaking to drop?

      • terranoid@lemmy.cafe
        link
        fedilink
        English
        arrow-up
        2
        ·
        4 天前

        I honestly would argue that we already made agi, and it’s just not as amazing as people thought it would be, and it doesn’t even mean self improving code in a meaningful way like we expected.

        LLM is called “narrow” AI but honestly I think they’re just trying to pretend this AGI is possible from scifi that isn’t. LLM tends to be okay at solving random problems it doesn’t have training with. Not amazing, but okay. It can solve general problems. I can describe a board game it never played and ask for advice and it would use other things it knows to solve it and it’s probably be fine. And maybe it doesn’t learn in the classic sense by improving model weights while working on stuff… But it already could rerun training automatically with new data and it’d just take a fuck long time.

        We have general learning AI and it kinda is okay at things, and I think that’s the massive mistake they made with AGI prophecies. We believed that if something could perform like humans and learn it could do it faster and with less resources. What if this is it and it kinda just doesn’t automatically improve itself meaningfully without infinite resources? Evolution worked on us for millions of years yet we don’t magically scale out infinitely. We have limits, and forget shit and make mistakes.

        I’m starting to think that they made massive assumptions about AGI based on extreme sci-fi ideas, finally created AGI, saw real world limitations, then said “well this can’t be agi if it hasn’t solved all our problems, keep working on it”. Yet, here it is, and it hallucinates and makes mistakes and won’t be that much better if you spend an extra trillion.

        We assumed improvement would be exponential and create ASI, but maybe it’s logarithmic and it’s fizzling out at “well it gets a lot of code right sometimes, and sometimes it doesn’t”.

        • HakFoo@lemmy.sdf.org
          link
          fedilink
          arrow-up
          2
          ·
          4 天前

          A legitimately interesting take.

          I can see two angles:

          1. AGI is implicitly expected to become ASI because many useful technologies have scaled to levels natural evolution didn’t.

          Our strongest weightlifters didn’t grow into thousand-horsepower cranes. Our quickest math savants are not cranking out a teraflop of FP32. But conversely, that’s why cranes and CPUs are economically interesting.

          1. Not really sure being able to carry on a plausible conversation on a “new subject” like the game example is much of an evidence of intelligence. ELIZA did it with a few kilobytes of RAM and some manually hardcoded rules. Having a huge set of training data just expands the opportunity for a plausible match.

          I think the lack of self-improvement and finite contexts are sort of damning. Outside of rare neurological disorders, I can’t make intelligent people forget how to ride a bike by reading the phone book to them.