Remix.run Logo
ckocagil 2 hours ago

Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.

hellohello2 10 minutes ago | parent | next [-]

Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.

irishcoffee 2 hours ago | parent | prev [-]

The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.