| ▲ | wronglebowski 3 hours ago | |
It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now. | ||
| ▲ | ComputerGuru 14 minutes ago | parent | next [-] | |
Just to play devil’s advocate: you can’t compare Qwen to a (proprietary/closed source) hosted model and deduce that Qwen is overthinking, as Qwen gives you the full reasoning/thinking trace while all the proprietary models now give you only a summary “to prevent distillation”, making it hard to properly compare apples to apples here. | ||
| ▲ | dannyw an hour ago | parent | prev [-] | |
Qwen thinking is really good in Mandarin; and probably natively trained the most there. Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language. | ||