Remix.run Logo
thevinter 4 hours ago

I'm ready to stand corrected, but I'm pretty positive that such a process would require 1) an insane amount of work and 2) wouldn't produce anything close to SOTA results because of the lack of training data.

It is my understanding that - sadly - the insane amount of copyrighted works and corporate crap is a prerequisite for having a corpus that is big enough

Schlagbohrer 4 hours ago | parent [-]

I think these days even the frontier labs are using large amounts of synthetic data too, which must be worth it even though it seems like an Ouroborous.