Remix.run Logo
irthomasthomas 5 hours ago

Why does anthropic change the set of benchmarks they use with every new model release?

https://www.anthropic.com/news/claude-opus-4-7

https://www.anthropic.com/news/claude-opus-4-6

pietz 5 hours ago | parent [-]

1. Benchmarks saturate 2. They select the most impressive improvments