| ▲ | 2001zhaozhao an hour ago | |
I would love to see models that can think at different rates and also output a thinking scratchpad alongside output text instead of before all output. Right now models need to rely on less legible compressed CoT to get high intelligence per token/step, but with diffusion they would just need to output more tokens per step instead. | ||