| ▲ | johnnyanmac 7 hours ago | |
In theory, sure. 1. Only use open source/CC compliant assets. 2. Acquire rights/licenses to any datasets that do not fit #1. e.g. the Google deal with Reddit for 60m/yr. 3. Offer programs to have creatives willingly submit their data, with some sort of residual output based on the number of times their assets are sampled. 4. If all that is still not enough, hire creatives to create assets for you. This is something Spotify did recently with "ghost artists"[0]. The intentions here are suspect, but a non-consumer facing artist providing work for an LLM wouldn't have the same ethical dilemmas 5. Lastly, if all that still isn't enough: governmental programs to either provide grants, subsidies, or more outreach to get the ball rolling. Would this cost tens, hundreds of billions of dollars? Yes. But clearly, that was not a barrier to entry for the industry anyway. So we can chalk this down to the personality of leadership or the wider culture of modern big tech [0]: https://harpers.org/archive/2025/01/the-ghosts-in-the-machin... | ||
| ▲ | fwip 5 hours ago | parent [-] | |
I agree with most of your post, but I don't think I agree that method 2 (Reddit deal) is necessarily ethical. I know it's too high of a bar for our nation to ever clear, but I think explicit author opt-in is the only ethical source of AI training data. Like, legally, I'm sure Reddit had the right to sell it, but probably over half their content was written before ChatGPT was ever announced. The TOS allowing reddit to make "derivative works" was largely understood to mean things like cropping photos, using your viral post in an ad, or maybe auto-translating your comment. | ||