| ▲ | thefxperson 2 hours ago | |
My understanding is that the extra token vectors generated as reasoning are still useful, but that their surface form (tokens themselves) do not necessarily reflect the underlying reasoning. i.e. reading the reasoning traces could be complete gibberish, but the hidden-dim vectors themselves still refine the latent probabilities and help in generating the correct answer. Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:
https://arxiv.org/html/2404.15758v1 | ||