| ▲ | ericd an hour ago | |
Gpt-oss—120b is like 1000 years old in AI years, whereas Qwen 3.8 27b is pretty young. What you’re seeing is that parameters aren’t apples to apples, and at a given parameter level, the new models are much, much better than the ones from a year or two ago. Like, to a comical degree. | ||
| ▲ | MichaelZuo 35 minutes ago | parent [-] | |
Wasnt this known by everyone who cared to pay attention? It practically became a joke about how a huge amount of the training data for GPT-4 was bottom of the barrel reddit vomit and obvious bot spam. Leading to many bizarre edge cases. | ||