| ▲ | stymaar 3 hours ago | |||||||
> insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!) It's a dense model so it will use all of its parameters per token. 37B active parameters isn't tiny at all, it's almost what Deepseek R1 had, and it's 2/3 of what Kimi k3 uses, so it's not going to be “insanely high” tps: it's going to be three times slower than Deepseek Flash (Prefil speed is going to be quite high though, but not token generation). | ||||||||
| ▲ | Azantys 2 hours ago | parent [-] | |||||||
Its 27B not 37B and having just 27B in total and 3T and like 30B active of those is still totally different. A 120B with 5B active is still much slower than a proper 5B. Just like the new Ling 3.0 Tiny with 8B and 1B active only gets around 120tk/s compared to 250tk/s which a real 1B one gets on my hardware. | ||||||||
| ||||||||