| ▲ | skerit 2 hours ago | |
> In Anthropic's testing, at its default "medium" effort the model matched or beat Claude Opus 5 at "high" effort on such tasks, in fewer steps and with fewer tokens Opus 5.5 has been amazing, but I'm confused by how this is worded. It "matched or beat" Opus 5? There is no matching. There is only surpassing. By miles. Like Opus 5 was the biggest disappointment of the year. Opus 5.5 is even better than Fable. I do not understand why they're not acknowledging it for the leap that it is? | ||
| ▲ | afro88 37 minutes ago | parent | next [-] | |
I don't get it either. Ditto for visual design capability. It's so far above Opus 5 and yet the announcement mentioned nothing about it. | ||
| ▲ | philipwhiuk an hour ago | parent | prev [-] | |
Underneath this means that you have say 50 tests and you grade each of them out of 10, then there was no test it did worse on. The data doesn't support it being better on every test (sometimes the score will be the same imperfect one, sometimes both will have gotten a perfect score). | ||