Remix.run Logo
gtirloni 21 hours ago

What's the relevance of the pelican benchmark when models probably saw it during training? Didn't OpenAI stop testing against SWE-Something because it was tainted?

simonw 15 hours ago | parent | next [-]

If they train for the benchmark, how come many of the pelicans produced by their different models at different reasoning levels still suck?

That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids.

genidoi 21 hours ago | parent | prev [-]

It's not a benchmark, it is a meme benchmark.

a3w 19 hours ago | parent [-]

Memes are arguably the web scale of benchmarks.

ljm 15 hours ago | parent [-]

AI reproducing Xtranormal video clips like NodeJS Is Web Scale should be the new benchmark.

If the dialogue is slop and not like the old memes then it fails.