| ▲ | simianwords 5 hours ago | |
They gesture at not using benchmarks for some reason... | ||
| ▲ | meric_ 4 hours ago | parent [-] | |
https://typesafe.ai/blog/antibenchmaxxing But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do | ||