| ▲ | embedding-shape a day ago |
| I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time. |
|
| ▲ | sunaookami a day ago | parent [-] |
| Same problem with every chinese model currently, they overthink way too much and take too much tokens and time. |
| |
| ▲ | embedding-shape a day ago | parent | next [-] | | More or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :) | |
| ▲ | EgregiousCube a day ago | parent | prev | next [-] | | A consequence of aggressive distillation? | |
| ▲ | szundi a day ago | parent | prev [-] | | [dead] |
|