| ▲ | breadislove an hour ago | |
On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report? | ||
| ▲ | krm01 30 minutes ago | parent [-] | |
Keeping track of any AI progress is becoming harder by the day, because there's ambiguity around common/clear/consistent benchmarks. Everything is constantly skewed into favourable directions. | ||