| ▲ | himata4113 2 hours ago | ||||||||||||||||
They have not increased in capabilities, they have increased in specialization. If you train a small model in another domain it will begin losing capabilities in the former domain. This is effectively the sigmoid problem. Although I will admit that if we discover a higher information density algorithm that it might change, but not by a substantial amount to where "super intelligence" in 1gb would be possible. | |||||||||||||||||
| ▲ | Philpax an hour ago | parent [-] | ||||||||||||||||
Over the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7... There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false. | |||||||||||||||||
| |||||||||||||||||