| ▲ | deaux 2 hours ago | |
"Intelligence" being what, math? Coding? Unfortunately there's a billion use cases for LLMs whose performance is not at all captured by the popular benchmarks they're all trying to maxx. | ||
| ▲ | whimsicalism 2 hours ago | parent [-] | |
if you are relying on a model for a business process, it should be simple enough to benchmark on that process | ||