| ▲ | chaboud 4 hours ago | |
Keep in mind that the model is thinking in a token space, itself a compressive representation of language. (Note: there's still a huge grammar penalty, so, ugh do think small.) | ||
| ▲ | qeternity 2 hours ago | parent | next [-] | |
The real breakthrough is going to be thinking in latent space. | ||
| ▲ | kzrdude 2 hours ago | parent | prev [-] | |
It selects tokens but they expand to embedding vectors which are huge, also in memory and attention requirements, I think? | ||