Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

  • 👁️👄👁️@lemm.ee
    link
    fedilink
    English
    arrow-up
    21
    arrow-down
    4
    ·
    1 year ago

    They aren’t reselling their information, they’re linking you to the source which then the website decides what to do with your traffic. Which they usually want your traffic, that’s the point of a public site.

    That’s like trying to say it’s bad to point to where a book store is so someone can buy from it. Whereas the LLM is stealing from that bookstore and selling it to you in a back alley.

        • BetaDoggo_@lemmy.world
          link
          fedilink
          English
          arrow-up
          8
          arrow-down
          1
          ·
          1 year ago

          So does any site that quotes the book. Just being trained on a work doesn’t give the model the ability to cite it word for word. For most of the books in this set you wouldn’t even be able to get a single accurate quote out of most models. The models gain the ability to cite passages from training on other sources citing these same passages.

          • BURN@lemmy.world
            link
            fedilink
            English
            arrow-up
            4
            ·
            1 year ago

            Being blocked by ChatGPT just means that the interaction layer you see doesn’t show the output, not that the output wasn’t generated.

            Everything you see that’s public facing and interfacing with an AI is an extreme filtering layer for what is output. There’s tons of checks that happen to ensure that they don’t output illegal content or any of a million other undesirable things.

          • 👁️👄👁️@lemm.ee
            link
            fedilink
            English
            arrow-up
            1
            ·
            1 year ago

            I’m too lazy and care too little but you can basically get it to roleplay as a book expert or something and to “remind” you of certain passages. It gets around the filter pretty easily, that’s how jailbreaks work.