Remix.run Logo
karmasimida 5 hours ago

Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5

Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.

_superposition_ 5 hours ago | parent | next [-]

I must be on the wrong X/Twitter then.

karmasimida 5 hours ago | parent [-]

Yes, please have it checked

avaer 2 hours ago | parent | prev | next [-]

I would trust 4chan more than I trust Twitter aura farming.

nsingh2 5 hours ago | parent | prev | next [-]

Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.

happycube 5 hours ago | parent [-]

Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else.

torginus 5 hours ago | parent | prev | next [-]

It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.

emp17344 5 hours ago | parent | prev | next [-]

Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.

ImprobableTruth 4 hours ago | parent [-]

Why would you accept it when the benchmark's ranking is obviously nonsense. It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.

dakolli 3 hours ago | parent | prev [-]

must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI.