• OfCourseNot@fedia.io
      link
      fedilink
      arrow-up
      6
      ·
      16 hours ago

      write out several thousand regex patterns

      Oh my! This thing is more likely to gain sentience than any llm.

      • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        3
        ·
        15 hours ago

        it is a truly absurd number, simply due to the inane BS variations on terminology and notation.

        It might have been more reasonable to train a local LLM on this material and then just query -that-… but hey 30ish seconds of claude’s free tier usage; and i’ve got a script that’s ~99.99% as good as a human in terms of pattern recognition. Pretty neat when you think about it.

        • OfCourseNot@fedia.io
          link
          fedilink
          arrow-up
          1
          ·
          9 hours ago

          And ‘truly absurd number’ might still be a bit of an understatement here. Saying that Claude can generate several thousand regexes for you is not really a selling point for llms in my book.

          On the other hand. Have you tried running them on the bible or something like that? I bet it finds it talks about ww2, 9-11, the nintendo wii, and shit…

          • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            1
            ·
            7 hours ago

            If, by bible, you refer to the holy book of the local Building Code; then no, I have not.

            and again; the regex in the script generated works, it’s just a teeny bit cursed and maybe a little bit sentient. Praise be to the Omnissaiah, and its blessed machine.

    • mobyduck648@lemmy.world
      link
      fedilink
      arrow-up
      6
      ·
      18 hours ago

      Can’t be worse than when I was an intern, the product was literally based around parsing arbitrary HTML with regex and not even decent regex, POSIX regex from like the '80s for reasons known only unto the programming gods.

    • ZILtoid1991@lemmy.worldOP
      link
      fedilink
      arrow-up
      3
      ·
      18 hours ago

      When I wrote my numerical parser, people told me to use regexes to detect what kind of numerical value is being parsed. I just did it in one go. Only regex-like checking is for ISO datetime values.

      • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        1
        ·
        15 hours ago

        regex is computationally pretty cheap, from what I’ve read; due to something about it being a simple AND gate or something… . IDR, and to be honest it doesn’t really matter in the era of 16-core/32thread 5GHz processors… but when you just want it to work… it does “just work”. /shrug.

    • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      1
      ·
      17 hours ago

      Not really. There will only ever be about a hundred people who ever use this tool, and it’s really just a massive timesaver that keeps people from needing to open a smorgasbord of files to get a few tidbits from each file; then collate it into one doc…

      If I get ~99.99% of all the information, and one hundred people use this, odds are, nobody will ever see the .01% edge case of missed info. and even if it does happen, it’s not like it’s obsfucated information in each file, it’s just a pain to open a hundred files and wait for them all to load at once through the disk-heavy GUI application we have to use.

      simpler to just read via the regex patterns, and then put it out into the terminal for the info needed.