| ▲ | vblanco 12 hours ago | |
Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k context size it means it wont do much before filling it context | ||
| ▲ | data-ottawa 10 hours ago | parent | next [-] | |
It’s definitely a heavy thinker, like most Qwens. I see strings like “write, now.” In the thinning traces then it goes on to think for a lot longer, so it’s kind of weird. I haven’t figured out hope to use this effectively yet on my strix halo. | ||
| ▲ | spwa4 11 hours ago | parent | prev [-] | |
We don't actually know how much thinking GPT and Opus do, the labs won't show us anymore. And they certainly take their time before starting to answer. | ||