Remix.run Logo
WASDx 3 hours ago

That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.

If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.