Remix.run Logo
hglaser 6 hours ago

Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.

Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

tomjakubowski 3 hours ago | parent | next [-]

Tasks are completed in about half the time too. Although we'll see if it slows down in a few weeks as Anthropic's model services are prone to do.

sharktheone 6 hours ago | parent | prev | next [-]

That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.

giancarlostoro 6 hours ago | parent [-]

Greed is not what's driving these prices, its cost. They considered very much in the red.

asdfasgasdgasdg 3 hours ago | parent [-]

The way to make money in this business right now is to make the absolute best product and convince everyone they need to use your thing, especially considering the training cost is a very large factor in the overall costs and you amortize that by selling inference.

user43928 5 hours ago | parent | prev | next [-]

Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5

onlyrealcuzzo 4 hours ago | parent [-]

The UI/UX seems impressively bad. DeepSWE's cost curve has a better, more obvious way to sort by only the top level of reasoning to avoid 80% of the graph just being the same 3-5 models at their 8 different reasoning levels...

It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.

I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.

An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.

makeavish 6 hours ago | parent | prev [-]

Nice catch, AA only shows max effort by default and I got disappointed thinking it's a token guzzler though: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level