| ▲ | cyanydeez 2 hours ago | |
I'm pretty sure the point of Qwen3.8-Flash-Next was to get the open source engines to integrate the qwen4 architecture. The fact that it basically broke open the local model supremacy was just a nice side effect. I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context. Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need. If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans. | ||