Remix.run Logo
adroitboss 3 hours ago

I'll tell you this. Output isn't too cheap to meter, there is no decoder.

krackers 3 hours ago | parent [-]

So an encoder-only model with a classifier trained on the heads or something? DeepSeek recently switched to an encoder-decoder architecture in an attempt to get the best of both worlds (fast prefill while preserving generation capability), I wonder if that might be the future?