Remix.run Logo
▲ zzzeek 7 hours ago

openai and anthropic trained on actually stolen data since it was pirated datasets.

google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.