Remix.run Logo
luciana1u 2 hours ago

i'll believe 'radically open' when the training data ships alongside the weights. until then it's a very fast demo.

adrian_b 2 hours ago | parent | next [-]

I just looked on Huggingface.co, and the training data is there.

For example, 3.3 Tbyte for code reasoning, 4.5 Tbyte for mathematical reasoning, 8.4 Tbyte of pre-train behaviors, and so on.

I did not compute the sum of the dataset sizes, but it appears to be some tens of Tbyte. Nonetheless, I assume that this amount of training data is more than an order of magnitude less than what OpenAI, Anthropic and the like have used, which must have been at least many hundreds of Tbyte, but more likely several thousands of Tbyte of data.

dakolli 2 hours ago | parent | prev [-]

Hey its a lot mpre thsn Anthropic which you probably use everyday all day without complaints.