Remix.run Logo
▲ PunchyHamster 7 hours ago

They can't be ethically sourced and good at the same time.

The current models intelligence depends on massive training dataset of essentially stolen data

▲mehrzad 7 hours ago | parent | next [-]

While that is true, theoretically a regulation could be enacted that output tokens must focus on STEM research and other practical tasks and the LLM must refuse tasks outside of those areas, just as Claude disallowed cybersecurity tasks. Obviously this would never happen, but the theft of the training data wouldn’t matter as much if the usecases were less sinister.

▲zzzeek 7 hours ago | parent | prev [-]

openai and anthropic trained on actually stolen data since it was pirated datasets.

google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.