| ▲ | embedding-shape 10 hours ago | |
Is the "Output Tokens per Intelligence Index Task" data actually correct or am I reading it wrong? It says there that "Kimi K3 (Max)" would think/reason less than than deepseek-v4-flash, and a whole bunch of other models, like less than hy3 and even gpt-oss-120b, but in my experience, K3 is probably the model that thinks/reasons the longest of all of these. Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point. | ||
| ▲ | Lwerewolf 4 hours ago | parent [-] | |
Just tried the preview on my little test codebase and a "check this out and tell me what you think" prompt used over double the tokens of the previous iteration, but it was a lot more eager as well. Kind of reminds me of the new laguna (s 2.1). | ||