Remix.run Logo
tdeck 2 hours ago

Why is the textual data bound up in these paper books worth the trouble?

These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?

For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.

awakeasleep an hour ago | parent | next [-]

I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!

rwmj an hour ago | parent | next [-]

Humans manage to be pretty intelligent with only reading perhaps a few thousand books in their lifetime. It seems unlikely that AGI will appear but only once it has read that last out of print 1983 book on knitting patterns.

tdeck an hour ago | parent | prev [-]

Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.

mapmeld an hour ago | parent | prev | next [-]

I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?

I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.

SpicyLemonZest an hour ago | parent | prev [-]

Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.