| ▲ | dennisy 5 hours ago |
| Are you able to share how it works in that case? |
|
| ▲ | adroitboss 4 hours ago | parent [-] |
| I'll tell you this. Output isn't too cheap to meter, there is no decoder. |
| |
| ▲ | krackers 4 hours ago | parent [-] | | So an encoder-only model with a classifier trained on the heads or something? DeepSeek recently switched to an encoder-decoder architecture in an attempt to get the best of both worlds (fast prefill while preserving generation capability), I wonder if that might be the future? |
|