| ▲ | Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s(github.com) |
| 90 points by snehesht an hour ago | 32 comments |
| |
|
| ▲ | kamranjon 5 minutes ago | parent | next [-] |
| Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well. https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen... |
|
| ▲ | deadbunny 34 minutes ago | parent | prev | next [-] |
| > Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository. And I thought piping to bash was bad |
| |
| ▲ | snehesht 33 minutes ago | parent | next [-] | | Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out. | |
| ▲ | gchamonlive 21 minutes ago | parent | prev [-] | | Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs. |
|
|
| ▲ | ryan_glass 2 minutes ago | parent | prev | next [-] |
| Anyone know how it compares to GLM 5.3 for real world use? |
|
| ▲ | snehesht an hour ago | parent | prev | next [-] |
| I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here. https://huggingface.co/Qwen/Qwen3.8-Flash-Next |
| |
| ▲ | roscas 4 minutes ago | parent | next [-] | | Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080. This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory. I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos. I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low. This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0. | |
| ▲ | thatsabadlook 22 minutes ago | parent | prev | next [-] | | Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me | |
| ▲ | proc0 an hour ago | parent | prev [-] | | Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions. | | |
| ▲ | incognito124 44 minutes ago | parent | next [-] | | Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore | | |
| ▲ | mickeyp 38 minutes ago | parent | next [-] | | I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result. It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop. | | | |
| ▲ | snehesht 41 minutes ago | parent | prev [-] | | Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course. | | |
| ▲ | nicce 24 minutes ago | parent [-] | | I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent. |
|
| |
| ▲ | thatsabadlook 20 minutes ago | parent | prev [-] | | Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model. | | |
| ▲ | geye1234 14 minutes ago | parent [-] | | I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that. |
|
|
|
|
| ▲ | Tepix 12 minutes ago | parent | prev | next [-] |
| Q2 quantization. Not interested. |
|
| ▲ | prettyblocks 28 minutes ago | parent | prev | next [-] |
| I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits). |
|
| ▲ | gdevenyi 42 minutes ago | parent | prev | next [-] |
| I had this working with the FreeToken inference engine a month ago when they launched. https://github.com/FlashML-org/FreeToken |
|
| ▲ | hypfer 28 minutes ago | parent | prev | next [-] |
| Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller? The Readme doesn't say, but it's all AI generated, so.. |
|
| ▲ | esafak an hour ago | parent | prev | next [-] |
| Has anyone calculated the effective intelligence of these quantized models? |
| |
| ▲ | nsagent 13 minutes ago | parent | next [-] | | See this recent paper: Quantization Degradation in Large Language
Models: A Signal–Noise Perspective [1]. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.[1]: https://arxiv.org/abs/2608.08188 | | |
| ▲ | merbanan 6 minutes ago | parent [-] | | I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation. |
| |
| ▲ | mkl 44 minutes ago | parent | prev [-] | | There's some info in the README, including: > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. https://github.com/Niko1221/Strata#which-model-should-i-pick | | |
| ▲ | nicce 27 minutes ago | parent | next [-] | | I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements? | |
| ▲ | javier2 40 minutes ago | parent | prev [-] | | ok that is getting interesting! |
|
|
|
| ▲ | panny 18 minutes ago | parent | prev | next [-] |
| I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM. |
| |
|
| ▲ | quietFalcon an hour ago | parent | prev | next [-] |
| Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context? |
| |
|
| ▲ | 0xbadcafebee 16 minutes ago | parent | prev [-] |
| Lol, sure, if you quant it to hell (Q2) it'll go real fast... |