Some of the world’s largest music publishers filed a blockbuster lawsuit against Anthropic late Friday night, alleging “one of the largest and most blatant ongoing thefts of intellectual property in history.”

  • danA
    link
    fedilink
    arrow-up
    1
    arrow-down
    1
    ·
    16 hours ago

    They destroy the books because it allows them to scan them bette

    That’s definitely one reason, but the copyright argument is also a part of it. The court explicitly said that their digitization is legal only if does not increase the number of copies of the book.

    There’s no need to do this legally - they could donate used books. The argument shows that the final product is transformative, so it doesn’t matter whether they keep the books or not.

    They have to keep the digital copy of the book because they add it to the training corpus for the LLM. Selling or donating the original physical book after doing that would void the fair use argument.

    LLMs can reproduce quite long passages of books, though I don’t think this changes the argument much, because it’s not reliable or useful.

    One of the tests that determines if it’s fair use or not is whether it can serve as a replacement for the original book. Pirated copies can, which is why they’re illegal. A summary like CliffNotes can’t. Even if the LLM can reproduce long passages, you can’t do that reliably (like you said) and it won’t work for all books.

    I always thought there was in obvious win where companies doing this could be forced to archive the scan publicly (after some period of time)

    I definitely agree with this. I think copyright law needs to be modernized to handle cases like this. I think the AI companies should be allowed to donate the digital copy to a library (like the Internet Archive) while still being allowed to keep their copy in their training corpus.

    • FishFace@piefed.social
      link
      fedilink
      English
      arrow-up
      2
      ·
      14 hours ago

      I think copyright law already needed to be modernised… and the way I think it could be done well is consistent with AI training needs. So the whole thing doesn’t really bother me. I’m surprised this angle incenses Lemmy so much because I’d have thought we’d on the whole by very anti-copyright in its current form.