Remix.run Logo
voidhorse 11 hours ago

Yeah, which is completely fine. There's a major difference between:

A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:

- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.

- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.

- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?

B: Company uses freely available scanned copy of the text:

None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.

I much prefer B.

Filligree 11 hours ago | parent [-]

B is illegal, and Anthropic ate a billion dollar fine for trying it, so you can’t even claim they don’t want to.

voidhorse 11 hours ago | parent [-]

Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?

Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.