| ▲ | gizmodo59 3 hours ago | |
I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The fact that muse and 3.8 were high a day back shows they are just the modern version of lmareana for the mass audience and PR stunts. | ||
| ▲ | jesuslop 3 hours ago | parent [-] | |
Is there something better over there you'd recommend? | ||