Remix.run Logo
ungut 2 days ago

Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place.

I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.

yorwba 2 days ago | parent [-]

I'm not aware of any successful accusations against OpenAI for illegally obtaining copyrighted material, in contrast to Anthropic, who settled for $3000 per work and then still had to buy legal copies to keep using them (likely for much less).

Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.

ungut 2 days ago | parent [-]

You are kind of right, but you also did not look very hard. They deleted huge datasets in anticipation of lawsuits, at least that much is known. Of course plaintiffs were unable to depose their internal lawyers (who apparently know why they were frantically deleted) due to 'attourney client priviledge' further refusing to provide any kind of transparency. But yeah, I guess they are they are better at covering their tracks and destroying evidence.

Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.

I think the real issues here are two-fold:

Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.

Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.