| ▲ | manofmanysmiles 3 hours ago | |
Imagine this, and sucesor models on Cerebras or other silicon... | ||
| ▲ | WASDx 3 hours ago | parent | next [-] | |
That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology. If tps is so high, a compaction step could be performed over every thinking turn to keep context size down. | ||
| ▲ | Moduke 2 hours ago | parent | prev [-] | |
Very exciting indeed. It is in the works. Their current dense offering, Gemma 4 31B, sits at ~1800t/s | ||