Remix.run Logo
wronglebowski 3 hours ago

It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.

ComputerGuru 14 minutes ago | parent | next [-]

Just to play devil’s advocate: you can’t compare Qwen to a (proprietary/closed source) hosted model and deduce that Qwen is overthinking, as Qwen gives you the full reasoning/thinking trace while all the proprietary models now give you only a summary “to prevent distillation”, making it hard to properly compare apples to apples here.

dannyw an hour ago | parent | prev [-]

Qwen thinking is really good in Mandarin; and probably natively trained the most there.

Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language.