| ▲ | kmeisthax 27 minutes ago | |
They didn't say that the Internet Archive wasn't "IP theft" - nor did they say that IP theft was inherently wrong. "AI training is infringement" is not exactly a copyright-maximalist view. The explicit training task used for pre-training is reproducing the content of the trained-on books; and models trained on such books are able to reproduce significant infringing chunks of them[0] unless specifically post-trained to refuse to do so. Additionally, they might have thought that Controlled Digital Lending was OK (it wasn't, but that's a different issue to AI training). As I've mentioned elsewhere in this thread, there's a common misconception that copyright is concerned with the number of copies in circulation as opposed to individual acts of copying. Or they don't care about any of that and just wanted to highlight the hypocrisy. [0] Which, under the "compression is intelligence" point of view, is entirely expected and not surprising in the slightest. | ||