| ▲ | PcChip 2 hours ago | |
Spelling mistakes? What inference engine are you using for flash next? | ||
| ▲ | anon373839 an hour ago | parent [-] | |
Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters. Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.) | ||