| ▲ | clhodapp a day ago | |||||||||||||
Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer. | ||||||||||||||
| ▲ | forgotTheLast 3 hours ago | parent | next [-] | |||||||||||||
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens. | ||||||||||||||
| ||||||||||||||
| ▲ | cyanydeez a day ago | parent | prev [-] | |||||||||||||
I assume theyre searching the local gradient to see if theres a better descent before proceeding. | ||||||||||||||
| ||||||||||||||