| ▲ | rcxdude 2 days ago | |||||||
It is and has been - skipping the tokens causes performance to drop. But replacing the tokens with filler causes the performance to drop, but by a lot less. You can also train models to emit broken or unrelated thinking tokens and their performance is also not much worse than the ones that are trained to output somewhat coherent thinking traces. This points to a hypothesis that the content of the tokens is only slightly related to the mechanism by which it improves performance, and that primarily the extra tokens allow the original prompt to be processed more deeply by the model, because earlier tokens will essentially pass through the model many more times than later ones. | ||||||||
| ▲ | neuroticnews25 a day ago | parent [-] | |||||||
I'll never understand some downvoters here. | ||||||||
| ||||||||