Remix.run Logo
Iolaum 4 hours ago

Yea and we are reaching the point where this benchmaxing is visible in the model's reported overthinking.

logicchains 4 hours ago | parent [-]

It's not overthinking, it's the right amount of thinking necessary for such a small model to get good results. The dumber the model, the more it has to think to be smart. There's no easy way to reduce the thinking without reducing the model quality.

zdragnar 4 hours ago | parent [-]

Qwen doom loops were amusing to watch the first time or two, but it's incredibly vexing to have it waffle over the same decision over and over and over and over again. I can get more done with a faster model by correcting it, and it feels better to babysit them than it does to babysit qwen to see if I need to intervene or if it will actually finish.

I do like the output from qwen when I get it, but honestly I haven't been impressed enough with it to put up with the downsides.

skohan 3 hours ago | parent [-]

It's only been a couple days, but I haven't seen looping issues with 3.8 so far, compared to 3.6 which did occasionally have this problem.