Remix.run Logo
▲ abejora an hour ago

Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

[1] Section 8.5 of the Sonnet 5.5 System Card

▲subscribed 4 minutes ago | parent | next [-]

I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

Anthropic made it that way, and I'd say the lower score is accurate.

▲eli an hour ago | parent | prev | next [-]

Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails

▲abejora an hour ago | parent [-]

You're right about its real world performance, and I worded my original comment wrongly.

I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

▲joeyhage 22 minutes ago | parent [-]

Claude, is that you?

▲radlad an hour ago | parent | prev | next [-]

I believe you meant to cite the Opus 5.5 System Card which states:

> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...

I cannot find a Sonnet 5.5 system card.

▲abejora an hour ago | parent [-]

It was linked in another HN post: https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c85...

▲Leary an hour ago | parent | prev | next [-]

And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!

▲oh_no 41 minutes ago | parent | prev | next [-]

it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

▲manojlds an hour ago | parent | prev | next [-]

Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?

▲MadameMinty an hour ago | parent | prev [-]

That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?

I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?

▲manojlds 44 minutes ago | parent [-]

Fallback was usually Opus 4.8