| ▲ | Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)(modelscope.cn) |
| 220 points by garo-pro 5 hours ago | 105 comments |
| |
|
| ▲ | SwellJoe 3 hours ago | parent | next [-] |
| Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio. I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context. And, MoE should make it run at a close to usable speed. |
| |
| ▲ | jubilanti 19 minutes ago | parent | next [-] | | A 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s The 3080ti is 912.4 GB/s | | | |
| ▲ | sosodev 2 hours ago | parent | prev | next [-] | | That’s only true if you think AI is the only reason to own a powerful and efficient server. Mine does plenty of traditional server stuff too. | | |
| ▲ | SwellJoe 2 hours ago | parent | next [-] | | I can do traditional server stuff on any old computer with a big hard disk and a decent amount of RAM. That's not worth $3500-$4000. When RAMpocalypse is over and we can buy a Strix Halo for under $2000 again, the math starts mathing. It becomes a pretty great desktop computer that also happens to run AI pretty well at a pretty good price. | | |
| ▲ | sosodev an hour ago | parent [-] | | Yeah, but that computer can’t also do the AI stuff. And not everybody has a desktop with multiple 32GB GPUs available. I’ll admit though I’m biased because I bought my board for $1600 back before the prices went crazy. | | |
| |
| ▲ | an hour ago | parent | prev | next [-] | | [deleted] | |
| ▲ | ArvidSu 2 hours ago | parent | prev [-] | | An "AI" server can do traditional server stuff but a traditional server can't do AI stuff (inference) |
| |
| ▲ | throwaw12 2 hours ago | parent | prev [-] | | how much performance (tok/s) can you expect from 128GB Strix Halo? assuming this model will be released with FP8 also can you use it for fine tuning? | | |
| ▲ | SwellJoe an hour ago | parent | next [-] | | The Strix Halo and DGX Spark are pretty danged slow, relatively speaking. I don't recall exact numbers, but with MoE models in this size ballpark (Laguna S 2.1), I seem to recall I was seeing about 20-25 t/s with a big context, which is close to usable. Qwen 3.8 27B crawls on this hardware, though, at 10-16 t/s, definitely not comfortable for interactive use. (Though this makes it seem like you can cook pretty good with a 4-bit ROCmFP4 quantization: https://github.com/julianmb/q38rocm the model does get notably dumber below six bits.) A model similar in size to Laguna S 2.1, but with only 6B active parameters, should be a notable amount faster, so I would imagine 25-30 t/s would be a reasonable guess for where Qwen 3.8 Flash Next will land. DFlash2 might improve all these numbers. It wasn't available last I was testing new models on the Strix Halo; I've only used MTP (which doesn't generally improve MoE models, but I believe DFlash2 can). Given software improvements, I'm hopeful an MoE in this size range will be the sweet spot that pushes past 40 t/s and is also smart enough for real work. Qwen 3.8 27B is finally a self-hostable model that's smart enough, but it thinks so hard it still isn't really useful for agentic interactive use. Note also prefill with large models is pretty slow on the Strix Halo (300 t/s, maybe). Time to first token is a painful wait, when using it interactively with large models. | |
| ▲ | downrightmike 7 minutes ago | parent | prev [-] | | You can only use up to 90gb for the GPU, so it doesn't fit | | |
|
|
|
| ▲ | ddtaylor 4 hours ago | parent | prev | next [-] |
| I enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful. OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win. However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame. I tried to get in contact with them at OpenRouter about this and I was interested in working with them in the past, but it's difficult to get in touch with the right people and they are growing very fast. I expect being acquired by Stripe will accelerate those problems in some ways. I have no doubt they will resolve all of these issues eventually and scaling that much that quickly is really hard, so kudos to them, but the road has been pretty lame and taken some wind out of my sails. |
| |
| ▲ | irthomasthomas 3 hours ago | parent | next [-] | | Openrouter was pretty great before prompt caching became common. Now it is extremely expensive for most individual workflows, unless you spend a lot of work customizing router preferences, and then you still get a worse cache hit rate than using the provider directly. I only keep $5-$10 in OR for occasional testing. | |
| ▲ | geek_at 3 hours ago | parent | prev | next [-] | | The best solution to this for me is to self host litellm or a different router and use model aliases. For example I have a model called "coding" and when a new good model comes out I just switch the backend without needing to change the alias or the key in my projects (opencode, etc). I have a few of them even a smart router called "agents" which will use local models but if it thinks the request might require higher reasoning it's routing to a different model | | |
| ▲ | try-working 3 hours ago | parent [-] | | I built a router that lets you route between local and cloud models. Link in my profile. | | |
| ▲ | embedding-shape 20 minutes ago | parent [-] | | Yeah, I also built my own "router" for this: if (process.env.LOCAL_MODEL {
http('localhost:3000/v1/completions')...
} else {
http('api.openrouter.ai/v1/completions')...
}
|
|
| |
| ▲ | npn 3 hours ago | parent | prev | next [-] | | I'm confused? Can you just define some presets and call them instead? With preset you can pinpoint a lot of things, especially the providers | |
| ▲ | ljlolel 3 hours ago | parent | prev [-] | | [dead] |
|
|
| ▲ | notnullorvoid 3 hours ago | parent | prev | next [-] |
| It will be interesting to see the intersection of this with inference engines like FreeToken which improve distribution of work for MoE models across CPU/RAM and GPU/VRAM. If all it takes for a competitive model to run locally at good speeds is a used 3090 and some DDR4, then we might be in for the year of local AI. https://github.com/FlashML-org/FreeToken |
| |
| ▲ | kamranjon an hour ago | parent | next [-] | | Have you tried FreeToken yourself? I was hoping to find some benchmarks on their github but took a quick pass at their research paper and it seems they're showing ~2x performance on qwen 3.6 35b when compared to llama.cpp - but llama.cpp is so sprawling and has so many options I find that a difficult comparison. | | |
| ▲ | notnullorvoid 16 minutes ago | parent [-] | | I haven't yet, though plan to when this model is released. The models that I've been daily driving (Gemma 4 26B, Qwen 3.8 27B) have fit nicely on my 3090. I think FreeToken only offers a perf increase for MoE models that you can't feasibly fit in VRAM. Yeah I'm sometimes unsure how to get best perf out of llama.cpp, and honestly thought it already did what the FreeToken paper discusses, but from everything I've been able to find since llama.cpp has no dynamic expert cache for GPU. An RFC discusses adding such capability and there's impressive results some are claiming from a fork, but I had to stop reading the thread, reading all the LLM generated comments and summaries from people was making me dizzy. RFC here https://github.com/ggml-org/llama.cpp/discussions/24528 which also links to some experimental implementations throughout the thread. |
| |
| ▲ | Zylokloto 3 hours ago | parent | prev [-] | | You can already run it locally its just not the same. It is still slow, a lot slower than what you are used to with claude and co. And as soon as you increase context size, your memory requirements jump. Then when it runs for 30 minutes for something claude needs 5, your device will get hot. And even a used 3090 is apparently now between 1-2k. | | |
| ▲ | notnullorvoid 2 hours ago | parent | next [-] | | > It is still slow, a lot slower than what you are used to with claude and co. That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus. In general we haven't reached the ceiling for what performance we can get out of consumer hardware. As evidence by FreeToken which hasn't even added MTP/speculative drafting support yet, which will add another boost. > Then when it runs for 30 minutes for something claude needs 5, your device will get hot. I doubt the timing differential here, but even still I run my 3090 pretty heavily with inference workloads and it stays cooler than when I use it for gaming. > And even a used 3090 is apparently now between 1-2k. Yeah I guess the price went up significantly in the last couple months, used to be hovering around 1k. 3090 isn't the only option though. | |
| ▲ | blahblaher 2 hours ago | parent | prev [-] | | yeah, but otoh... f* Anthropic and OpenAI |
|
|
|
| ▲ | syntaxing 2 hours ago | parent | prev | next [-] |
| Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls. |
| |
| ▲ | corysama 16 minutes ago | parent | next [-] | | So, I know https://cactuscompute.com/needle is designed only to enable tool calling on tiny devices. But, I wonder if anyone has used it as a CPU-side mediator between a tool and a GPU-side local LLM making semi-natural-language tool requests... | |
| ▲ | cpburns2009 2 hours ago | parent | prev [-] | | Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today. | | |
| ▲ | SparkyMcUnicorn an hour ago | parent | next [-] | | Have you given Ornith-1.5-35B a shot? It's been a pretty decent step up for me compared to Qwen3.6 https://news.ycombinator.com/item?id=49362401 | | | |
| ▲ | 2 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | puzzlingcaptcha 2 hours ago | parent | prev | next [-] | | What sort of pp/tg speed do you get on a Strix Halo? | | |
| ▲ | cpburns2009 an hour ago | parent | next [-] | | This is the best I got, all with Unsloth's quantizations. Laguna-S-2.1:UD-Q4_K_XL (no MTP)
pp=186.4 t/s
tg=27.8 t/s Qwen3.6-35B:UD-Q4_K_XL (with MTP)
pp=404.4 t/s
tg=83.2 t/s Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP)
pp=343 t/s
tg=12.1 t/s Laguna actually performed better than I remembered. I thought it was slower. | | | |
| ▲ | ascii0eks84 10 minutes ago | parent | prev [-] | | What are pp/tg? I get 30t/s on 27B qwen. | | |
| ▲ | throwawayffffas a minute ago | parent [-] | | pp is prompt processing how fast it processes the prompt. Tg is token generation how fast, it generates tokens. |
|
| |
| ▲ | cyanydeez 2 hours ago | parent | prev [-] | | it'll hopefully improve with more MoE and half the prefill/generation. I think it's the sweet spot for the strix halo for smarter or vibe tasks. |
|
|
|
| ▲ | fcanesin 4 hours ago | parent | prev | next [-] |
| HF link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next |
|
| ▲ | big-chungus4 4 hours ago | parent | prev | next [-] |
| > We are releasing these architectural improvements ahead of time so that the community can prepare for the upcoming full family of Qwen4 models. That gives me hope that "full family" means it will include smaller models like 4B. |
|
| ▲ | pwython 4 hours ago | parent | prev | next [-] |
| I was already rolling around the idea of a 128GB M5 Max MBP. Now this! A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range. |
| |
| ▲ | Eric_WVGG 2 hours ago | parent | next [-] | | Just out of curiosity, why run "local-local" when you could just set up a Mini or Studio at home and query it over http? [edit] whole conversation about this in another thread https://news.ycombinator.com/item?id=49433413 I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out. | | |
| ▲ | LeBit an hour ago | parent [-] | | This is the way. I’m doing that. Mac Mini M4 Pro with 48G RAM as a headless llama.cpp server. I much prefer using " thin clients " as the interface to the big VMs running in my homelab |
| |
| ▲ | sscaryterry 4 hours ago | parent | prev | next [-] | | I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something... | | |
| ▲ | smcleod 3 hours ago | parent [-] | | 50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B? | | |
| ▲ | Casteil 3 hours ago | parent | next [-] | | I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65. | | |
| ▲ | syntaxing 2 hours ago | parent [-] | | With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further |
| |
| ▲ | sscaryterry 2 hours ago | parent | prev [-] | | I tried 8-bit, perhaps I should try 6-bit. |
|
| |
| ▲ | irthomasthomas 3 hours ago | parent | prev | next [-] | | IDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs? | | |
| ▲ | smcleod 3 hours ago | parent [-] | | That was mainly before the M4 generation when they didn't have matmul instructions. | | |
| ▲ | jasonjmcghee 3 hours ago | parent [-] | | M5 prefill is much faster than M4. I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max. For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max. For larger dense models, some fraction of that, but similar multiple. | | |
| ▲ | smcleod 3 hours ago | parent [-] | | Yes, I have the M5 Max. But there was no matmul acceleration before the M4 which made things a lot slower. |
|
|
| |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
|
|
| ▲ | c16 2 hours ago | parent | prev | next [-] |
| +1 to the long list of people hoping for Qwen3.8-27b A3B. |
| |
| ▲ | NitpickLawyer an hour ago | parent | next [-] | | They've said no moe for 3.8, and since they're already releasing a qwen4 early preview, they're probably focusing on that arch going forward. | |
| ▲ | WiSaGaN an hour ago | parent | prev [-] | | You probably meant Qwen3.8-35B-A3B. But judging from some of the words from their team, it seems unlikely unfortunately. | | |
| ▲ | vorticalbox 7 minutes ago | parent [-] | | They normally release a 35b dense and an 27b moe (4B active per token) For context 35B on my m4 runs at 10 tokens a second, 27B moe runs 50-60 tokens a second. |
|
|
|
| ▲ | Catloafdev 2 hours ago | parent | prev | next [-] |
| Very curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse. |
| |
| ▲ | natrys 2 hours ago | parent | next [-] | | Why not? It's not really competing in the same size class. Besides, as they explicitly wrote here, the main goal for this release is not performance, rather to serve as a reference for inference runtimes about what to implement. So that later Qwen 4 can be released with zero day support. | | |
| ▲ | Catloafdev an hour ago | parent [-] | | Good point, I didn't see that. I guess I categorized them in the same bucket of 'runs on 128gb machines' Guess Qwen 4 is the one to wait for. |
| |
| ▲ | NitpickLawyer an hour ago | parent | prev [-] | | Their "next" variants are usually undercooked, but useful for the community to verify support for inference stacks. This will likely be the same. |
|
|
| ▲ | hedora 3 hours ago | parent | prev | next [-] |
| Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days). Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory. I’m wondering if it can replace claude for llm-friendly coding tasks. |
| |
| ▲ | cpburns2009 3 hours ago | parent [-] | | So back in the Qwen 3.5 release, the 122B-A10B model scored slightly better than the 27B model. I'd expect this new 125B-A6B to perform similarly to the recently released 27B. Qwen3.8 27B is supposed to rival Sonnet/Opus 4.6. | | |
| ▲ | hugmynutus 2 hours ago | parent | next [-] | | Qwen3.8/Qwen3.6 has a weird self doubt/thinking too much problem. You can prompt it away. I would say it "approximates" Opus 4.X class models well enough especially for coding/linux problems. The only reason I stopped using it as much is I was getting 25-35tok/s on Intel B70 (non-quant) which made some responses slow. For a long running/autonomous task, it would probably be sufficient. | |
| ▲ | wongarsu 2 hours ago | parent | prev | next [-] | | There is the rule of thumb that if you take the geometric mean of the total and active parameters of an MoE model you get the equivalent size of an equally capable dense model. If you follow that formula, you would expect a 125b-a6b model to match a 27b model (sqrt(125*6) = 27.3). That does not feel like a coincidence | |
| ▲ | hedora 3 hours ago | parent | prev | next [-] | | Thanks. My current stack ranking of anthropic models is: 4.6 ~= 4.8 4.7 much worse. Fable and newer consistently tells me to pound sand, so I’m not sure what I’m paying $200/month for. 4.8 sometimes does too, but it’s at least usable most of the time. So, I’d expect this to mostly replace Claude for my workflows. The main tradeoff for me should mostly be token throughput vs. no longer really trusting anthropic. | |
| ▲ | cyanydeez 2 hours ago | parent | prev [-] | | I've got the A10B hooked up to deer-flow and it does remarkable well when you dont need to baby sit it. |
|
|
|
| ▲ | honestlyranked 4 hours ago | parent | prev | next [-] |
| Alibaba is giving sleepless nights to the tech giants |
| |
| ▲ | drannex 9 minutes ago | parent | next [-] | | To be fair, Alibaba IS a tech giant, one of the biggest in fact. They are just giving sleepless nights to the western tech giants. | |
| ▲ | WithinReason 2 hours ago | parent | prev [-] | | Sounds like a line from a fairy tale |
|
|
| ▲ | fkndkfn an hour ago | parent | prev | next [-] |
| I can feel Dario Amodei's tears in the announcement :) |
|
| ▲ | system2 13 minutes ago | parent | prev | next [-] |
| These companies are naming their products worse than I was naming my half-baked software in the 90s as a junior developer. |
|
| ▲ | big-chungus4 4 hours ago | parent | prev | next [-] |
| I hope there is going to be a free endpoint... Unlike 35B-A3B, I am nowhere close to running it locally |
|
| ▲ | freddiehdxd 41 minutes ago | parent | prev | next [-] |
| Does it support vision? |
|
| ▲ | bellowsgulch 3 hours ago | parent | prev | next [-] |
| Really happy for those with 128GB+ RAM. Sitting here with my Apple M1 Max with 64GB though. Was looking forward to a Qwen3.8-35B-A3B like many others. |
| |
| ▲ | dofm 3 hours ago | parent [-] | | Have you tested Muse Glimmer in low reasoning strength? Token generation is slow (and prefill is) but you will likely find it solves actual problems faster than Qwen 3.6 35B-A3B. | | |
| ▲ | bellowsgulch 3 hours ago | parent [-] | | I’ll give it a try! Thanks for the heads up! | | |
| ▲ | dofm 2 hours ago | parent [-] | | I’m using the Unsloth 4-bit quant. To change the reasoning strength you just put text in the system prompt. From memory it is: Reasoning strength: low
|
|
|
|
|
| ▲ | lousken 2 hours ago | parent | prev | next [-] |
| gpt oss killer? this can easily run on a server cpu with its memory bandwidth |
| |
| ▲ | Art9681 an hour ago | parent | next [-] | | You must have hibernated for a year. Most modern 27b models can outperform gpt-oss-120b. | | |
| ▲ | bearjaws 28 minutes ago | parent [-] | | Not exactly surprising given it's a dense model at 2.7x the size of the experts in gpt-oss |
| |
| ▲ | stymaar an hour ago | parent | prev [-] | | gpt-oss is long dead though. It's been 8 month since Qwen 3.5 was released. |
|
|
| ▲ | dmead 2 hours ago | parent | prev | next [-] |
| This is great. I have a weird system layout (192gb system ram, 8gb vram). the mixture of experts models have been nice when i can run the dense reasoning layers on the gpu (which somehow fit?!) and then the expert on the cpu. its worked out to to 40 tokens/seconds on their 80b-a3b model. we'll see how much of a hit this is. |
|
| ▲ | cogman10 4 hours ago | parent | prev | next [-] |
| Wow. I wasn't expecting this. I thought they were going to do a 35B model instead. |
| |
| ▲ | hasteg 3 hours ago | parent [-] | | As a 5090 owner and local model enthusiast, I was hoping it would be 35B A3B so I could run it myself =(. | | |
| ▲ | Tuna-Fish 3 hours ago | parent | next [-] | | The 27B one is great on a 5090. This one is basically aimed at macs, Strix halo and DGX Spark. | |
| ▲ | cpburns2009 3 hours ago | parent | prev [-] | | You can run the 27B released last week. I haven't tried it yet myself but the 3.6 version runs great on my 5090. | | |
|
|
|
| ▲ | tarruda 5 hours ago | parent | prev | next [-] |
| Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page. |
| |
| ▲ | petu 5 hours ago | parent | next [-] | | It was in description under the countdown initially, but was quickly removed. It also said 51B of n-grams and new attention (IIRC it said "Qwen Sparse Attention"). edit: here's a random screenshot https://x.com/AiBattle_/status/2092210011858460819/photo/1 | |
| ▲ | NitpickLawyer an hour ago | parent | prev [-] | | This is what I copied from the en version of the modelscope page, right when they published it: > Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token. > Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork. There was another paragraph about a new attention, but I didn't copy that. |
|
|
| ▲ | BrucecarlL 4 hours ago | parent | prev | next [-] |
| Waiting for the performance report! Ai hope it can beat DS |
|
| ▲ | isatty 3 hours ago | parent | prev | next [-] |
| Can I run a fp8 quant with 96gb VRAM? |
| |
| ▲ | cpburns2009 3 hours ago | parent [-] | | Only VRAM? Unlikely unless you can also load the whole model into regular RAM. The previous 3.5 release was 250gb at BF16, so FP8 would likely be around 125gb. Your best best is FP4/Q4. | | |
|
|
| ▲ | Alien1Being an hour ago | parent | prev | next [-] |
| AI VENDOR PRESS RELEASE |
|
| ▲ | tw1984 4 hours ago | parent | prev | next [-] |
| Qwen4 sounds exciting |
| |
|
| ▲ | cyanydeez 2 hours ago | parent | prev | next [-] |
| oooh, I like a6b; that will be nice. 3.5 A10B qwen works really well in deer-flow when you want to seriously vibe code or research and you're just not going to baby sit. |
|
| ▲ | onesandofgrain 2 hours ago | parent | prev | next [-] |
| where are the humans geez |
|
| ▲ | blurbleblurble 4 hours ago | parent | prev | next [-] |
| gg |
| |
| ▲ | david927 3 hours ago | parent [-] | | Well put and succinctly put. And if OxA is a flash model? it becomes: goodnight |
|
|
| ▲ | mrdoe 4 hours ago | parent | prev | next [-] |
| lol blocked with dns4eu what a joke this resolver has become |
|
| ▲ | metrofun 3 hours ago | parent | prev [-] |
| [dead] |