| ▲ | karmasimida 5 hours ago | |||||||
Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5 Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion. | ||||||||
| ▲ | _superposition_ 5 hours ago | parent | next [-] | |||||||
I must be on the wrong X/Twitter then. | ||||||||
| ||||||||
| ▲ | avaer 2 hours ago | parent | prev | next [-] | |||||||
I would trust 4chan more than I trust Twitter aura farming. | ||||||||
| ▲ | nsingh2 5 hours ago | parent | prev | next [-] | |||||||
Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index. | ||||||||
| ||||||||
| ▲ | torginus 5 hours ago | parent | prev | next [-] | |||||||
It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability. | ||||||||
| ▲ | emp17344 5 hours ago | parent | prev | next [-] | |||||||
Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark. | ||||||||
| ||||||||
| ▲ | dakolli 3 hours ago | parent | prev [-] | |||||||
must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI. | ||||||||