| ▲ | ls612 10 hours ago | |
On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too. | ||
| ▲ | spijdar 10 hours ago | parent | next [-] | |
I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality. That's just vibes, though. | ||
| ▲ | KaoruAoiShiho 10 hours ago | parent | prev [-] | |
https://www.reddit.com/r/LocalLLaMA/comments/1wj3s31/thank_y... | ||