| ▲ | toofy a day ago | ||||||||||||||||||||||||||||||||||||||||||||||||||||
> But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest. only if they share the scans/transcripts and don’t hide it away. | |||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | DougBTX a day ago | parent [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||
Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book! | |||||||||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||||||||