| ▲ | ac29 an hour ago | |
gpt-oss-120b only has 5B active parameters, so its not surprising Qwen3.8 27B outperforms it (Qwen3.8 is also ~13 months newer, which is forever in LLMs) | ||
| ▲ | kennywinker an hour ago | parent | next [-] | |
Fair enough. I’ve barley touched oss-120b, so i didn’t know it was so few active params. For a direct comparison, qwen3.6-35b-a3b is still better at coding than oss-120b. And Qwen3.8-27b is still better at coding than opus 4.1. Yes, if you list off models 27b is better than it’s all older models. But that’s my point - newer models are better than older models at the same AND much smaller size. That’s because model size matters less than they say. Training data and model architecture matter more. | ||
| ▲ | anon373839 an hour ago | parent | prev [-] | |
No, it’s not the active parameters. Qwen 3.8 Flash has 6B active and it smokes both models. | ||