Remix.run Logo
selectodude 13 hours ago

When you realize that LLMs are “just” extremely efficient lossy data compression, it’s hard for me to see how it’s anything other than taking people’s shit, putting it into a gigantic zip file, and letting people search against it.

gruez 12 hours ago | parent [-]

Wait till you hear about Perfect 10, Inc. v. Amazon.com, Inc. (2007) and Authors Guild, Inc. v. Google, Inc. (2015), both of which ruled that lossy and verbatim copies (respectively) are allowed for for-profit use.

selectodude 12 hours ago | parent [-]

Too late. Authors Guild, Inc. v. Google, Inc. is a good one too because Internet Archive got the exact opposite outcome in court for doing the exact same thing. I recognize the bullshit, I just call it out to keep myself sane.

gruez 12 hours ago | parent [-]

>Internet Archive got the exact opposite outcome in court for doing the exact same thing

No, it's not the same thing. Contrary to what many people think, "fair use" isn't something you can invoke to do whatever copyright infringement you want. The judge is supposed to consider several factors, one of which is whether the work was "transformative". In google's case it was offering search results. Internet archive was operating a "digital library" (aka. a filesharing site). Whatever you hate about AI companies sucking up electricity and displacing jobs, they're certainly more transformative (and arguably more transformative than even google search) than whatever the internet archive was doing.

selectodude 12 hours ago | parent [-]

That’s not true. 1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own. 2. Google provided access to the whole book, that’s why they got sued.

If I run a book through AES, that’s pretty transformative too!

gruez 12 hours ago | parent [-]

>1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own

And has this been tested in court? After all, you see people uploading tv shows on youtube, then pasting a snippet of fair use in the description. That doesn't make it true. If anything, the unfavorable ruling for internet archive suggests libraries were incorrect with their interpretation of the law.

>2. Google provided access to the whole book, that’s why they got sued.

No it didn't. From wikipedia:

"For works still under copyright, Google scanned and entered the whole work into their searchable database, but only provided "snippet views" of the scanned pages in search results to users."