Remix.run Logo
▲ cyanydeez 2 hours ago

I'm pretty sure the point of Qwen3.8-Flash-Next was to get the open source engines to integrate the qwen4 architecture.

The fact that it basically broke open the local model supremacy was just a nice side effect.

I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context.

Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need.

If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans.