| ▲ | thevinter 4 hours ago | |
I'm ready to stand corrected, but I'm pretty positive that such a process would require 1) an insane amount of work and 2) wouldn't produce anything close to SOTA results because of the lack of training data. It is my understanding that - sadly - the insane amount of copyrighted works and corporate crap is a prerequisite for having a corpus that is big enough | ||
| ▲ | Schlagbohrer 4 hours ago | parent [-] | |
I think these days even the frontier labs are using large amounts of synthetic data too, which must be worth it even though it seems like an Ouroborous. | ||