| ▲ | mdp2021 2 days ago | |
> For example, even if you make thinking tokens literally just Generally speaking yes, but actually no (just randomness is suboptimal, adding steps just to add steps is suboptimal). There is a mechanism working there (in having a CoT) that is not quite clear. The task is to optimize the efficiency of CoT. Understanding that it is not a plain "chain of thought" is the start of the problem, the solution is not there yet. If we had the solution, there would exist no overthinking - CoT would be optimal (lean and essential plus best results). | ||