| ▲ | amelius 4 hours ago | |||||||
What I want is a model that is trained with data that is openly available, where the data is curated by academia. I don't want corporate crap in my AI (unless it has been filtered properly). | ||||||||
| ▲ | thevinter 4 hours ago | parent | next [-] | |||||||
I'm ready to stand corrected, but I'm pretty positive that such a process would require 1) an insane amount of work and 2) wouldn't produce anything close to SOTA results because of the lack of training data. It is my understanding that - sadly - the insane amount of copyrighted works and corporate crap is a prerequisite for having a corpus that is big enough | ||||||||
| ||||||||
| ▲ | dorkypunk 3 hours ago | parent | prev [-] | |||||||
There are models that do that, for example the Olmo family of models, although they have Gemma 3 performance levels for that matter. | ||||||||