| ▲ | gtirloni 21 hours ago | ||||||||||||||||
What's the relevance of the pelican benchmark when models probably saw it during training? Didn't OpenAI stop testing against SWE-Something because it was tainted? | |||||||||||||||||
| ▲ | simonw 15 hours ago | parent | next [-] | ||||||||||||||||
If they train for the benchmark, how come many of the pelicans produced by their different models at different reasoning levels still suck? That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids. | |||||||||||||||||
| ▲ | genidoi 21 hours ago | parent | prev [-] | ||||||||||||||||
It's not a benchmark, it is a meme benchmark. | |||||||||||||||||
| |||||||||||||||||