| ▲ | driverdan 7 hours ago |
| Judge Alsup issued the original order that determined they were liable for piracy but that training LLMs on books was fair use. It's worth reading if you're interested in the topic. https://www.courtlistener.com/docket/69058235/231/bartz-v-an... |
|
| ▲ | tzs 5 hours ago | parent | next [-] |
| Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber. He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1]. [1] https://www.theverge.com/2017/10/19/16503076/oracle-vs-googl... |
| |
| ▲ | dataflow 2 hours ago | parent [-] | | He learned Java to understand the Oracle v. Google case better. His middle name is Haskell. |
|
|
| ▲ | penguin_booze an hour ago | parent | prev | next [-] |
| So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use? It's funny that crimes can be settled in cash. IOW, everything has a price; and the price is always right. Settlement ought to be the euphemism for blood money. In addition to the settlement, what I'd consider fair is to have these companies pay royalties in perpetuity. Of course, that's not tractable. |
| |
| ▲ | shakna 30 minutes ago | parent | next [-] | | > So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use? No, that's what they got in trouble for - a lack of consent. If the author consents, it would have been fine. If they bought the books, then it is fine. Digitisation through destruction, like most book scanning systems. As long as the original work is destroyed during the process, and you actually paid for it, then it is fair use. If it regurgitates, then the author can sue you again. So you are incentivised to make damn sure it doesn't. That's not covered by fair use. Its only if the original cannot be accessed anymore, and you paid to get the original. Both must be true, for fair use to hold. | |
| ▲ | owenfi an hour ago | parent | prev | next [-] | | Yeah, I feel like penalties here should be something like 10% of revenue in perpetuity. Then companies might think twice about asking forgiveness instead of permission. | |
| ▲ | sdenton4 26 minutes ago | parent | prev | next [-] | | I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain? No, because that's not what copyright is about. Learning from and building on previous work is civilization. Copyright maximalism is a plague. | |
| ▲ | simianwords 31 minutes ago | parent | prev [-] | | Why do they need prior consent? What sort of rent seeking do you want? |
|
|
| ▲ | brlewis 4 hours ago | parent | prev | next [-] |
| Yes, that is interesting. It sounds like he was aware of the theoretical possibility of a book being regurgitated verbatim. Do you know if he was aware it had been done? https://news.ycombinator.com/item?id=49000742 If he was not aware, I wonder if he still would have described the process as "exceedingly transformative" had he been aware. |
| |
| ▲ | FeepingCreature 2 hours ago | parent [-] | | Note that they're testing for 100-word passages. This is a level of memorization that avid readers can credibly also reach. Note also that Sonnet 3.7 had to be jailbroken. Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization. | | |
| ▲ | globular-toast 2 hours ago | parent [-] | | No "avid reader" could recall anywhere near that much text. That takes dedicated effort to commit to memory. Copyright was never meant to stop people copying books anyway, it was meant to stop machines (ie. printing presses) copying them. Edit: Apologies, I misread it as "100 pages". My point about copyright still stands, though. | | |
| ▲ | FeepingCreature an hour ago | parent | next [-] | | I disagree that avid readers cannot complete entire passages from books they've read several times when fed a prefix. | |
| ▲ | JAlexoid an hour ago | parent | prev [-] | | I used to use the initial letters of a whole paragraph from Lord of The Rings as my password. Some of us have a good enough memory. |
|
|
|
|
| ▲ | 3 hours ago | parent | prev [-] |
| [deleted] |