| ▲ | ByteAtATime an hour ago | |
Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do | ||
| ▲ | anewhnaccount2 8 minutes ago | parent | next [-] | |
Yes this is the reason. Transformations like quantisation preserve the rank of outcomes much better than probability mass. | ||
| ▲ | RussianCow 33 minutes ago | parent | prev [-] | |
I think it also helps that speed and latency aren't as vital for coding and we can afford to let the model think for longer. | ||