Remix.run Logo
glimshe 19 hours ago

The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."

Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?

eru 19 hours ago | parent | next [-]

> Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time?

No. They also use lots of other methods to get training data.

npn 18 hours ago | parent | prev | next [-]

No but with 100% clean data you can easily train a model to filter ai generated content.

ACCount37 14 hours ago | parent | prev | next [-]

As a rule: all high quality text is useful.

There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.

At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.

Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.

TiredOfLife 11 hours ago | parent | prev [-]

Also there was a highly discussed paper talking about how “touched by machines” content will kill llms. About a month after the papers first llms trained with “touched by machines” content appeared an the capabilities of the models got huge upgrade by using that dirty content