Remix.run Logo
▲ RussianCow 2 hours ago

It's fast. The average speed is 115 tokens/sec according to OpenRouter. I haven't tested the model to see how it is in practice, but I'd certainly pay a little extra for faster inference.

Edit: Though the average latency of 1.5s isn't very low, so it might not be that fast in practice for agentic work. Also, I don't know how much thinking it does, as that's generally been the drawback to Chinese models.

▲celrod 2 hours ago | parent [-]

Using Artificial Analysis

Model | Reasoning | Intelligence Index | Artificial Analysis million output tokens for the intelligence index -|-|-|- Step 5 | ? | 44 | 160 GLM 5.3 | Max | 45 | 210 MiMo V2.6 Pro | ? | 46 | 140 Kimi K3 | Max | 44 | 160 Qwen Max 0902 | ? | 45 | 190 DeepSeek 4.1 Flash | Max | 39 | 250 GLM 5.3-flash | Max | 42 | 180 GPT-6 Astra | Low | 46 | 10

It looks reasonable by open model standards. This doesn't capture the fact that DeepSeek and Step 5 have much higher token/s than the rest, other than MiMo Ultraspeed. MiMo V2.6 is either slow but cheap, or fast but expensive. Just based on these numbers, it looks good. Astra-low is one of the fastest and cheapest because it doesn't use many tokens, but I've never tried it. I liked DeepSeek and GLM when I used them.