• Goodlucksil@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    22
    arrow-down
    2
    ·
    7 days ago

    Before anyone goes on a wild goose chase, this excerpt is on the opening of Chapter 4 instead of 5 in the first edition.

    Anyone surprised to find that public domain art is determined to be AI generated does not realize that feeding an LLM requires a lot of art, and public domain is the easiest one to come across. AI detectors, as far as I know, check AI signs, and (most) AI signs came from their use by public domain artworks.

    • prole@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      4
      ·
      6 days ago

      Anyone surprised to find that public domain art is determined to be AI generated does not realize that feeding an LLM requires a lot of art, and public domain is the easiest one to come across.

      Yeah no we definitely realize that.

      Wouldn’t that mean that it should be better at identifying public domain work?

    • melsaskca@lemmy.ca
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      6 days ago

      So, all the AI bullshit floating around, along with the requisite tech bros owners, is beneficial to society as a whole because they steal art that was no longer protected by copyright? My first question is always “If it all went away would anything change?”.

  • InvalidName2@lemmy.zip
    link
    fedilink
    English
    arrow-up
    27
    arrow-down
    1
    ·
    7 days ago

    I’ve been commenting on the open, public internet in some form or fashion since the 1900s. These models have been trained to sound like me, not the other way around. Fortunately, the bullies at the back of the class who don’t pay attention have mostly caught on to this fact, after having it repeated to them hundreds of times. The other day I was thinking about it, and realized it’s been a bit since anybody’s accused me of using AI to write a comment. I think it’s because I stopped proofreading what I write. Leave in those mistakes and the assholes at the back of the room who don’t pay attention are too naive to think that machines can’t make mistakes, I guess.

  • Wildmimic@anarchist.nexus
    link
    fedilink
    English
    arrow-up
    16
    ·
    7 days ago

    I postulated exactly these issues with detecting LLM code in software projects, making “no LLM code”-type policies unenforceable. It doesn’t work most of the time now, and it will only get worse in the future.

    • Tenderizer@aussie.zone
      link
      fedilink
      English
      arrow-up
      1
      ·
      6 days ago

      If nothing else, banning AI code forces them to hide it. Which is itself in some ways an advantage.

  • HugeNerd@lemmy.ca
    link
    fedilink
    English
    arrow-up
    11
    ·
    6 days ago

    Shouldn’t it be easy with today’s computing power to simply search all content first for a match to known sources?

    • Zink@programming.dev
      link
      fedilink
      English
      arrow-up
      8
      ·
      6 days ago

      That would make sense if you wanted to do what was best for the user rather than push the Chosen Product.

    • Jerkface (any/all)@lemmy.ca
      link
      fedilink
      English
      arrow-up
      4
      ·
      6 days ago

      Wouldn’t necessarily tell you if that known source is AI. So now we have to go through everything by hand anyway.

      • HugeNerd@lemmy.ca
        link
        fedilink
        English
        arrow-up
        3
        ·
        6 days ago

        Of course it can, if it’s from 1859 the chances of it being AI are … I leave that as an exercise for the reader.

        • FatCrab@slrpnk.net
          link
          fedilink
          English
          arrow-up
          5
          ·
          6 days ago

          And how do you know the source for your “hit” isn’t lying? You are not describing a trivially easy problem at all.

  • LePoisson@lemmy.world
    link
    fedilink
    English
    arrow-up
    11
    ·
    7 days ago

    I made a machine with the purpose of literally creating output that plausibly sounds like a human wrote it.

    Now that it creates things that sound plausibly like they came from a real person it’s a problem. Now quick make something else that pretends to know it’s not human when the whole point of the damn machine is to sound human.

    Tjciiktn it’s ridiculous. HMsie?$+'izhav!&(:8_! This world is so fucked

    • TwodogsFighting@lemdro.id
      link
      fedilink
      English
      arrow-up
      7
      ·
      6 days ago

      There was an old lady who swallowed a fly, I don’t know why she swallowed a fly – perhaps she’ll die!

      There was an old lady who swallowed a spider That wriggled and jiggled and tickled inside her; She swallowed the spider to catch the fly; I don’t know why she swallowed a fly – perhaps she’ll die!

      There was an old lady who swallowed a bird; How absurd to swallow a bird! She swallowed the bird to catch the spider That wriggled and jiggled and tickled inside her, She swallowed the spider to catch the fly; I don’t know why she swallowed a fly – perhaps she’ll die!

      There was an old lady who swallowed a cat; Well, fancy that, she swallowed a cat! She swallowed the cat to catch the bird, She swallowed the bird to catch the spider That wriggled and jiggled and tickled inside her, She swallowed the spider to catch the fly; I don’t know why she swallowed a fly – perhaps she’ll die!

      There was an old lady that swallowed a dog; What a hog to swallow a dog! She swallowed the dog to catch the cat, She swallowed the cat to catch the bird, She swallowed the bird to catch the spider That wriggled and jiggled and tickled inside her, She swallowed the spider to catch the fly; I don’t know why she swallowed a fly – perhaps she’ll die!

      There was an old lady who swallowed a goat; Just opened her throat and swallowed a goat! She swallowed the goat to catch the dog, She swallowed the dog to catch the cat, She swallowed the cat to catch the bird, She swallowed the bird to catch the spider That wriggled and jiggled and tickled inside her, She swallowed the spider to catch the fly; I don’t know why she swallowed a fly – perhaps she’ll die!

      There was an old lady who swallowed a cow; I don’t know how she swallowed a cow! She swallowed the cow to catch the goat, She swallowed the goat to catch the dog, She swallowed the dog to catch the cat, She swallowed the cat to catch the bird, She swallowed the bird to catch the spider That wriggled and jiggled and tickled inside her, She swallowed the spider to catch the fly; I don’t know why she swallowed a fly – perhaps she’ll die!

  • adr1an@programming.dev
    link
    fedilink
    English
    arrow-up
    11
    ·
    7 days ago

    false positives exist, and the real monster is that it also happens with “faces” of criminals.

  • hirihit640@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    7
    arrow-down
    2
    ·
    7 days ago

    Yeah yeah AI system has false positives. Like a human, really. In fact humans are pretty bad at detecting AI text: https://arxiv.org/html/2211.13087

    Among AI models, Blenderbot stood out: in AI-AI conversations, it was judged human 67% of the times - more often than actual human-human conversations. In human-AI conversations, human participants were labeled as humans 68% of the time, and AIs were classified as AI 56% of the time.

    (FYI, a detection rate of 50% is the same as random chance, so 56% is pretty poor)

    • Don_alForno@feddit.org
      link
      fedilink
      English
      arrow-up
      11
      ·
      7 days ago

      The point of the post is there is no way to reliably detect it as it just reproduces what it was trained on.

      A human also did not take a trillion $ of investment to be able to produce this false output.

      • hirihit640@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        6
        ·
        6 days ago

        The point of the post is there is no way to reliably detect it as it just reproduces what it was trained on.

        I hope people realize this before they call out other comments as “AI generated”, something I see all too often on Lemmy

        • Don_alForno@feddit.org
          link
          fedilink
          English
          arrow-up
          1
          ·
          5 days ago

          First, we’re the humans. Investments into the existence and wellbeing of humans don’t have to be justified by an ROI. It’s kind of all we’ve been doing since we came out of the caves. A tool on the other hand needs to justify it’s existence by saving us more time and money than it cost to make or enabling us to do things we couldn’t do before, otherwise it’s a shitty tool.

          Second, the LLM that claimed Frankenstein was slop took tens or hundreds of billions to make. And that cost doesn’t scale with the demand. To be able to run even one instance of it you have to spend the entire lump sum of training the model. On the other hand you can train one more human if you need one more skilled worker and you don’t have to redevelop the entire curriculum from scratch to do that.

  • redditStinksSuperBad@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    arrow-down
    9
    ·
    edit-2
    7 days ago

    It’s not proven but some folks think she wasn’t ahead of her time and stole her story. I wouldn’t believe it myself but it’s sort of hard when she was hanging around “Castle Frankenstein” that was rumored to have a grave robbing scientist that lived there who did experiments to bring creatures back to life…

    From a bot: The breakdown of the claims behind this theory includes:The “Mad Scientist” of Castle Frankenstein: Johann Conrad Dippel, an alchemist and physician, was born at Burg Frankenstein (near Darmstadt, Germany) in 1673. Local legends—largely popularized long after Shelley’s time—claimed he practiced grave robbing, boiled bones to create an “elixir of life” (known as Dippel’s Oil), and attempted to transfer souls between corpses.

    The Travel Connection: In 1814, Mary Shelley, her step-sister Claire Clairmont, and Percy Bysshe Shelley traveled down the Rhine River. They stopped in Gernsheim, Germany, which is roughly 10 miles away from Castle Frankenstein. This proximity led some pop-historians to theorize that she visited the castle or heard local tales of Dippel from boat captains along the river.

    The Lack of Documentation: Mary Shelley’s detailed journals and letters from that 1814 trip make no mention of Castle Frankenstein, Dippel, or local ghost stories. When the travelers ran low on money, they rushed back to England, leaving very little time for unauthorized detours.

    Where the Name Came From: Shelley famously stated that the core concept of the novel came to her during a waking dream at Lord Byron’s house in Switzerland during the “Year Without a Summer” (1816), after a late-night ghost-story challenge. As for the name “Frankenstein,” it is a real German surname meaning “stone of the Franks”. Linguistic and literary historians note that it was a striking, gothic-sounding word she likely encountered or naturally synthesized, without needing a specific real-world blueprint. Ultimately, while the parallels between Dippel’s legends and Victor Frankenstein’s fictional misadventures make for a fascinating modern urban legend,

      • tmyakal@infosec.pub
        link
        fedilink
        English
        arrow-up
        6
        arrow-down
        1
        ·
        7 days ago

        Just a run-of-the-mill misogynist who can’t accept the idea that a woman wrote a piece of foundational genre fiction. See: all the unfounded claims that actually her husband wrote the story.

        • redditStinksSuperBad@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          2
          ·
          7 days ago

          Relax. I do think she wrote the book herself but also took heavy inspiration from Dippel and never credited the story. This would be like someone hanging around Castle Dracula, writing a book called Dracula, and then claiming the story “came to me in a dream”.

  • mrmisses@lemmy.world
    link
    fedilink
    English
    arrow-up
    140
    arrow-down
    2
    ·
    7 days ago

    Does it think it’s AI because it was trained on it so it recognizes it as being part of AI now?

    • glimse@lemmy.world
      link
      fedilink
      English
      arrow-up
      45
      arrow-down
      38
      ·
      7 days ago

      It was posted like some kind of gotcha but…this book was in the dataset.

      I’m aware that LLMs don’t keep the dataset in memory but it “knows” that this is an existing work but these sites aren’t doing any wild calculations to figure out if it was AI-generated. They basically just check to see if the sentences exist elsewhere and they do. In the original dataset.

      So it was wrong to say it’s AI-generated but it correctly identified that it’s not original.

      • michaelmrose@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        2
        ·
        7 days ago

        Everything you said was hallucinated. Nothing you said was even slightly coherent or connected with reality. Absolutely nothing was correct. If you let your cat dance on the keyboard more sense would come out.

        • glimse@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          ·
          6 days ago

          That must have sounded really clever and powerful in your head for you to hit send, huh? Couldn’t decide which zinger to go with so you sent all 4?

          I’d have given you more credit if you just posted the quote from Billy Madison. Try harder lol

          • michaelmrose@lemmy.world
            link
            fedilink
            English
            arrow-up
            1
            arrow-down
            1
            ·
            6 days ago

            No it just annoys me when people have no understanding of how shit works and just make up an explanation instead of looking it up. Hell you could have asked chatGPT and probably got a correct answer.

      • Grimy@lemmy.world
        link
        fedilink
        English
        arrow-up
        95
        ·
        7 days ago

        They analyze statistical patterns, they don’t cross reference the training data.

        It’s picking up the text as generated because it’s mostly guess work and constantly spits out false positives, especially with non native speakers.

        • ferret@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          11
          ·
          7 days ago

          Honestly I didn’t expect them to flag non-native speakers. LLMs don’t often make gramatical mistakes (or at least not ones your average joe could identify)

          • SirLeToet@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            arrow-down
            1
            ·
            7 days ago

            Weirdly enough i saw a massive increase in little typoes on reddit this year. More typoes and such in a few months than i saw in years.

            It feels like LLMs started making mistakes on purpose.

            • Axolotl@feddit.it
              link
              fedilink
              English
              arrow-up
              11
              ·
              7 days ago

              And are also less prone to use slang unless they go more deep in the culture

            • M137@lemmy.today
              link
              fedilink
              English
              arrow-up
              7
              ·
              7 days ago

              Oooh yeah, the most common dumb mistakes are (by my own checking) done by Americans in the VAST majority cases. Stuff like not knowing how to use your/you’re, there/their/they’re, cloths/clothes, definitely/defiantly, where/were/we’re, its/it’s etc. correctly and the many other failings of basic grammar like “would/could of” and so much else.
              I don’t think there’s another country in the world where the native speakers are so fucking bad at even the most basic shit.
              The most common response I’ve gotten and seen when others correct stuff like that is “calm down, English probability isn’t their main language” yet when checking it’s been Americans who made the mistakes 99% of the time.

              • Enkrod@feddit.org
                link
                fedilink
                English
                arrow-up
                5
                ·
                7 days ago

                I was going to say:

                The difference is between growing up, learning a language by hearing and making the sounds, and learning it as a foreign language, having to understand and memorize the rules.
                “They’re, their, there” all sound the same as well as “could’ve, could of”.

                But interestingly other English speakers don’t make those mistakes as often as people emerging from the american education system.

                • wols@lemmy.zip
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  ·
                  6 days ago

                  It’s plausible that Americans are worse at this than foreigners, but English spelling is a crime and everyone who makes mistakes ought be excused on account of how ridiculous the “rules” are.
                  I know how to spell most of the words I use and I resent the fact that my memory is polluted with this otherwise useless knowledge.
                  In many other languages if you know how to say a word, you automatically know how to spell it. We need a world language with grammar as simple as English and spelling as straightforward as German.

      • FishFace@piefed.social
        link
        fedilink
        English
        arrow-up
        33
        arrow-down
        1
        ·
        7 days ago

        That is not in the least bit how a tool like this works.

        All AI detection is extremely unreliable, but they operate on principles which, if the assumptions supporting them were true, would be sound. The way you imagine they work is different: you’re describing a “novel text detector” which is not at all the same thing as an “AI text detector”.

        AI detection works, at a very high level, by throwing a bunch of examples of AI text, and a bunch of examples of non-AI text, into a machine-learning model, and training it until it is able to recognise the AI examples as such. It doesn’t work by throwing in all existing text including novels written before AI, because that would never do what you want.

        (The reason, if you’re interested, why this ends up not working is because the features such a model identifies as indicative of AI text are very sensitive: if you train it on Claude and ChatGPT, it will not work on Gemini output. If you train it on Gemini, when Gemini updates it will get worse. If someone generates text with a weird prompt, it may slip by. If someone writes in a weird way, it may get flagged. And if any AI company wanted to defeat AI detectors, it could trivially feed its output through one during the training process and give that output as examples to avoid in the training, so that the model learns to create output which doesn’t “look like” AI output to those detectors.)

        • glimse@lemmy.world
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          2
          ·
          7 days ago

          I was being overly simplistic - I meant more that the patterns it’s trying to detect were created by an LLM trained on the data they’re inputting.

          It’s like how reddit comments from 2016 look generated. If you stick one of those into an AI detector, it’ll give a false positive for the same reason

          • michaelmrose@lemmy.world
            link
            fedilink
            English
            arrow-up
            4
            arrow-down
            1
            ·
            6 days ago

            You weren’t simplifying what you said was just wrong what the person you are responding to is just correct. Just admit when you are wrong

        • glimse@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          4
          ·
          7 days ago

          It’s original in the book. What the user entered is an exact copy.

          “Original” is not the correct work but you know what I mean.

            • glimse@lemmy.world
              link
              fedilink
              English
              arrow-up
              3
              arrow-down
              5
              ·
              7 days ago

              Yes…because the pasted text has some of the exact patterns that are part of the dataset that the website trained on…because that dataset contains text generated from a dataset containing the exact paragraph…

              I feel like you’re being deliberately obtuse.

              • [deleted]@piefed.world
                link
                fedilink
                English
                arrow-up
                7
                arrow-down
                1
                ·
                7 days ago

                I feel like you don’t understand the difference between ‘AI generated’ and ‘something that existed over 200 years ago’.

                • glimse@lemmy.world
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  arrow-down
                  5
                  ·
                  7 days ago

                  I feel like you don’t understand that LLMs train on text that existed 200 years ago…

          • RampantParanoia2365@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            7 days ago

            Yes, but if that were true, then every single book ever published would be flagged. The training data is for teaching it patterns, not just cross-referencing.