Remix.run Logo
▲ ramish94 2 hours ago

In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.

Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)

FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)

CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)

Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family

▲level87 2 hours ago | parent | next [-]

This is crazy, what is the point of all these equivalent models?

▲eli 2 hours ago | parent | next [-]

Those are just 3 particular technical benchmarks. Presumably Opus is a larger model and has greater world knowledge.

▲salviati 2 hours ago | parent | prev [-]

Price going down on each release

▲bbor 2 hours ago | parent | prev | next [-]

Yup. Recursive self improvement presented in hard numbers.

▲bigyabai 2 hours ago | parent | prev [-]

It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.

OpenAI and Anthropic's lead is vanishingly small at this point.

▲TuxSH 2 hours ago | parent | next [-]

> OpenAI and Anthropic's lead is vanishingly small at this point.

Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.

Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.

▲SubiculumCode 2 hours ago | parent | prev | next [-]

Yeah, I did kind of feel like the step down from Opus 5.5 was so large as to never make it appealing.

▲bbor 2 hours ago | parent | prev [-]

Your takeaway from "Sonnet 5.5 matches and sometimes exceeds the SoTA worldwide" is "their lead is vanishingly small"...?