Remix.run Logo
finn888 15 hours ago

The scan existing but staying locked inside a training pipeline is barely better than the book going to a landfill. At least make the raw scans available.

jeroenhd 13 hours ago | parent [-]

Lossily stuffing books into a model through a training process has so far been deemed legal. Enabling piracy by giving away digital copies is not. The internet archive tried to give away digital copies and they got sued to hell and back (which they should've seen coming from miles away).

I'm sure they have some repository available somewhere. They can even sell the digital copies down the line if they're done with them (only once, of course).

Joel_Mckay 13 hours ago | parent [-]

>Lossily stuffing books into a model through a training process has so far been deemed legal

Not in the EU, UK, or US. "AI" companies were already forced to settle their piracy cases, but often they get a free pass by law enforcement via regulatory capture.

The problem is a book author contracted publisher does not assign legal rights of duplication to a company/individual that buys a legitimate print. It can take over 70 years in some places to become public domain.

The core issue is "AI" firms have so much borrowed cash around, that getting a $1.5B fine for being a pirate is taken as a cost of doing business. The law is simply not equipped to handle this type of hyper-scaling criminal act. =3

https://www.bbc.co.uk/news/articles/c5y4jpg922qo

jeroenhd 11 hours ago | parent [-]

The piracy lawsuits have so far only deemed that obtaining books through pirating is a violation of copyright. Training the models on them and reselling model access has yet to be ruled illegal, from my understanding, even though there has been plenty of opportunity to.

Anthropic's crime wasn't stealing the contents of books and making a derivative work of it, but torrenting a shitload of books. Had they bought all the ebooks, I don't think the lawsuit would've gone anywhere.

Joel_Mckay 11 hours ago | parent [-]

>I don't think the lawsuit would've gone anywhere

That is not how copyright/trademark/contract laws treat similar works. Most LLM know about Disney Mickey Mouse, and LLM vector search space proximity results will gravitate more accurate reproductions of protected works regardless of granularity of data.

OpenAI simply canceled a popular service to avoid Disney wrath.

https://www.theglobeandmail.com/world/article-openai-sora-di...

Also, trying to escape directly ripping off notable famous people with nonunion talent:

https://www.youtube.com/watch?v=YhgYMH6n004

I would say the "AI" firms will keep buying time with all that borrowed cash. Everything that could be scraped has already been stolen, and thus the problem should begin to self-correct. The Shrek movie release market correction history correlation is funny, and a new film is due 2027 in July. =3