Remix.run Logo
sourcecodeplz 13 hours ago

it is not just copyrighted works.

openai and anthropic and google etc hire PHDs to create datasets by solving problems, writing thinking traces.

then they train their models on those datasets.

also as another user said, what is this 80+ years to protect some text? its ridiculous

classified 11 hours ago | parent [-]

It's only ridiculous for the big thieves, but if you violate copyright, you'll pay dearly. Calling it ridiculous will not protect you.

therealpygon 10 hours ago | parent [-]

Thieves…like Anthropic who literally just settled a case for pirating? Those kind of thieves?

Always better to name them outright.

Anthropic STOLE data for training. Verifiably. Admitted. In court. Then settled.[0]

You can also likely assume that OpenAI did too. There simply wasn’t enough data without having done so at least at some point in their corporate history. In fact, I seem to recall they were able to extract Harry Potter text from all the major modes[1]

https://fortune.com/2026/07/21/anthropic-copyright-settlemen... [0]

https://arxiv.org/abs/2601.02671 [1]