Remix.run Logo
yorwba 2 days ago

AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742

And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.

spwa4 2 days ago | parent [-]

But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.

yorwba 2 days ago | parent [-]

GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection.

spwa4 2 days ago | parent [-]

You mean very likely the Anna's archive torrent dump because it's MUCH better quality than the general internet and beyond a certain amount of input data (which is a lot, but much less than the internet) the only thing that matters in training is the quality of the data, to the point that now many labs have thousands of people just making and improving essentially school exercises full time?

Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.

yorwba a day ago | parent [-]

No, I mean the WebText corpus whose construction from 45 million Reddit post with at least 3 karma is described in section 2.1 of the PDF I linked. They did remove all Wikipedia documents.

Anna's Archive didn't exist in 2019.