| ▲ | tormeh a day ago | |||||||||||||||||||||||||||||||
Is qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious. | ||||||||||||||||||||||||||||||||
| ▲ | ch_sm a day ago | parent | next [-] | |||||||||||||||||||||||||||||||
In my experience, yes. A bit more reliable than gemma for me. I mostly use A3B (35B, mix of experts) though, because it‘s faster, and in the same ballpark intelligence wise as the dense 27B, so it’s the sweetspot for me. I want to try cohere‘s mini code model next, but worried the runtimes aren‘t optimized for that yet. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | androiddrew a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
I have been running 3.6 27b on a dual AMD r9700 setup using Opencode and Matt Pocock's skills workflow for writing Golang CLIs. It's decent, but won't win any awards on code architecture. I guess you can try to AGENTS.md the deficits but I am just exploring its raw Opencode experience right now. Much slower than an API but still 3x times faster than I can read. Tuning it in with a community chat template and a specific penalty for repeats was the sauce needed to get it to work. I can probably start loop daddying it now over the tickets Matt's flow creates. So yeah, it's the best local model I've seen. I am going to try the Qwopus 3.6 fine tune soon with the same spec and tickets and compare the output of both. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | seanmcdirmid a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
I actually have long discussions with Gemini about this and have wound up download a bunch of different models for different things. There is no best, just fast but worse, slow but better, agentic or not, reasoning or not great at large contexts, better world knowledge, uncensored, etc…. It’s a bit daunting actually since there isn’t really a one size fits all model that you can just use for everything. | ||||||||||||||||||||||||||||||||
| ▲ | ernsheong a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Yes it's between this and Gemma 4 31B which is much slower, but looks like it won't ever get an upgrade. I have to conclude that the MoE variants are unreliable, and MTP sometimes just can't get tricky formatting right. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | hnfong a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
People have been able to run DeepSeek v4 flash with a high spec Mac. | ||||||||||||||||||||||||||||||||
| ▲ | schaefer a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
I flip flop between qwen 3.6 27b and qwen 3.6 35b 4b active. But there’s also the quantization of DeepSeek v4 flash called dwarfstar | ||||||||||||||||||||||||||||||||
| ▲ | cmrdporcupine a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
Gemma4 models are arguably better. Or at least about the same. | ||||||||||||||||||||||||||||||||
| ▲ | atemerev a day ago | parent | prev [-] | |||||||||||||||||||||||||||||||
The best model you can run locally is Kimi K3, as long as you have the hardware. If "what model I can still run on a something resembling something I can put on desktop without separate electricity and cooling water inputs", then it is probably GLM 5.2 (can be run on e.g. Nvidia DGX Station workstation). As long as you have about $100k-$150k. | ||||||||||||||||||||||||||||||||