| ▲ | girvo 3 hours ago | |||||||
The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too. | ||||||||
| ▲ | jonsoft an hour ago | parent | next [-] | |||||||
I made this 3D game in a day on the same setup with Qwen Code as agent: https://games.jonathanpage.com/ And I am not a web developer! It's an extraordinary model. (Mouse and keyboard required) | ||||||||
| ▲ | walrus01 3 hours ago | parent | prev [-] | |||||||
Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it. | ||||||||
| ||||||||