| ▲ | tarruda 6 hours ago | |
Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page. | ||
| ▲ | petu 6 hours ago | parent | next [-] | |
It was in description under the countdown initially, but was quickly removed. It also said 51B of n-grams and new attention (IIRC it said "Qwen Sparse Attention"). edit: here's a random screenshot https://x.com/AiBattle_/status/2092210011858460819/photo/1 | ||
| ▲ | NitpickLawyer 2 hours ago | parent | prev [-] | |
This is what I copied from the en version of the modelscope page, right when they published it: > Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token. > Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork. There was another paragraph about a new attention, but I didn't copy that. | ||