Remix.run Logo
amelius 4 hours ago

What I want is a model that is trained with data that is openly available, where the data is curated by academia. I don't want corporate crap in my AI (unless it has been filtered properly).

thevinter 4 hours ago | parent | next [-]

I'm ready to stand corrected, but I'm pretty positive that such a process would require 1) an insane amount of work and 2) wouldn't produce anything close to SOTA results because of the lack of training data.

It is my understanding that - sadly - the insane amount of copyrighted works and corporate crap is a prerequisite for having a corpus that is big enough

Schlagbohrer 4 hours ago | parent [-]

I think these days even the frontier labs are using large amounts of synthetic data too, which must be worth it even though it seems like an Ouroborous.

dorkypunk 3 hours ago | parent | prev [-]

There are models that do that, for example the Olmo family of models, although they have Gemma 3 performance levels for that matter.