Remix.run Logo
0cf8612b2e1e 4 hours ago

Do you expect any of the labs to have an accompanying data dump with: here’s every book ever written, newspaper article, song lyric, Disney movie, GitHub repo, etc. Oh, and we obviously never paid for any of this.

Even if you did, I doubt training is bit-for-bit reproducible, so you will always have to take someone’s word for the final artifact.

anon373839 15 minutes ago | parent [-]

These OSS purity spirals are not helpful. Fully open source LLMs are great resources, but they underperform for the reason you stated: they can't use the same training data. Surely the people making these complaints know this. So I find myself wondering: are they just really pedantic? Or is this issue getting astroturfed?