| ▲ | simonw 15 hours ago | |
If they train for the benchmark, how come many of the pelicans produced by their different models at different reasoning levels still suck? That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids. | ||