Remix.run Logo
▲ snehesht 2 hours ago

I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

▲roscas an hour ago | parent | next [-]

Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

▲StumpChunkman 34 minutes ago | parent [-]

How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.

▲roscas 7 minutes ago | parent [-]

Yes, 3080 with 10GB, forgot to mention that.

Mine is at the moment writting some cpp code for some SBOM tests.

I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

Oh I will run some other tests with hermes now because hermes is amazing too.

▲proc0 2 hours ago | parent | prev | next [-]

Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.

▲incognito124 2 hours ago | parent | next [-]

Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore

▲mickeyp 2 hours ago | parent | next [-]

I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

▲snehesht an hour ago | parent [-]

This is interesting, thanks. - https://github.com/Neroued/ninfer

▲snehesht 2 hours ago | parent | prev [-]

Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.

▲nicce an hour ago | parent [-]

I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.

▲thatsabadlook an hour ago | parent | prev [-]

Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.

▲geye1234 an hour ago | parent [-]

I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.

▲PcChip 44 minutes ago | parent [-]

Spelling mistakes?

What inference engine are you using for flash next?

▲anon373839 14 minutes ago | parent [-]

Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

▲thatsabadlook an hour ago | parent | prev [-]

Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me