| ▲ | CMay 2 hours ago | |
> if you run one model, run glm-5.3 That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing. Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case. Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set. | ||