Remix.run Logo
manofmanysmiles 3 hours ago

Imagine this, and sucesor models on Cerebras or other silicon...

WASDx 3 hours ago | parent | next [-]

That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.

If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.

Moduke 2 hours ago | parent | prev [-]

Very exciting indeed. It is in the works. Their current dense offering, Gemma 4 31B, sits at ~1800t/s

https://news.ycombinator.com/item?id=49308715