Remix.run Logo
Squarex 2 hours ago

I don't know why, but the benchmarks still fails to cover the difference between large models and small ones. The small ones are great for many things, including general coding, but the larger ones, like fable and astra, have some kind of intelligence that is not present in the small ones.

sinuhe69 an hour ago | parent [-]

More parameters = more facts stored. Knowledges are almost incompressible, where strong reasoning only requires a 3B core or so.

yorwba 44 minutes ago | parent [-]

Weibo's VibeThinker manages with half of that: https://arxiv.org/abs/2511.06221 (They finetuned Qwen2.5-Math-1.5B for reasoning.)