| ▲ | jonsoft 2 hours ago | |||||||
It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache): | ||||||||
| ▲ | girvo an hour ago | parent [-] | |||||||
Ah that’s a shame. GLM 5.3 Flash is honestly as good IMO and can run on two pretty successfully from what I understand. I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware | ||||||||
| ||||||||