Remix.run Logo
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard(artificialanalysis.ai)
59 points by aarondong 4 hours ago | 33 comments
andy99 31 minutes ago | parent | next [-]

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.

afavour 19 minutes ago | parent | next [-]

What are you asking that you’re so regularly running into censorship?

wewtyflakes 4 minutes ago | parent | next [-]

I've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.

wild_egg 6 minutes ago | parent | prev | next [-]

I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.

Retr0id a minute ago | parent [-]

I haven't been using it for long, but so far the refusals seem about on par with how things were on Opus 4.8.

icedrift 14 minutes ago | parent | prev [-]

If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in

jefftk 11 minutes ago | parent [-]

I thought we were talking about Opus 5, the model Fable now falls back to?

buzzerbetrayed 17 minutes ago | parent | prev | next [-]

Yep. I cancelled my Claude Max subscription 2 weeks ago after feeling like Anthropic was doing everything it could to fuck with my day to day. Their lead would have to become significant for me to ever go back.

pinkyboy 13 minutes ago | parent | prev [-]

Fascinating that one of your biggest problems is entirely hypothetical to your experience.

chmod775 31 minutes ago | parent | prev | next [-]

The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.

At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

ricardobeat 2 minutes ago | parent [-]

The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

firasd 42 minutes ago | parent | prev | next [-]

Very interesting that one of the components is "AA-Omniscience Index"

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.

This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)

I've thought for a while that Gemini 3.x has 'big model smell'

zormino 9 minutes ago | parent | prev | next [-]

I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.

aarondong 4 hours ago | parent | prev | next [-]

Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...

eli an hour ago | parent | next [-]

Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.

emmp 9 minutes ago | parent [-]

Indeed, you can filter the graphs to see these the values for alternative reasoning settings of the models. Opus 5 High reasoning scored 59 on the index (exactly the same as GPT 5.6 Sol Max), and costs $1.06 per task (vs $1.04 Sol Max). So these seem essentially equivalent on both metrics.

midnightbobarun 4 hours ago | parent | prev [-]

5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too

nijave an hour ago | parent | next [-]

I think on swebench verified luna was only like 3% points lower for 1/5 the cost

Like 96% vs 93% or something

impulser_ an hour ago | parent | prev | next [-]

It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.

scrlk 41 minutes ago | parent | next [-]

Not just compute for OAI, GPT-5.6 is more token efficient across the board vs the Anthropic equivalents: https://artificialanalysis.ai/models?intelligence-index-toke...

No wonder why Tibo can afford to hit the reset button liberally.

charcircuit 32 minutes ago | parent | prev [-]

I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

giancarlostoro an hour ago | parent | prev | next [-]

Probably because they made ASICs to run inference for less.

brookst an hour ago | parent [-]

Are those actually deployed at scale yet?

brcmthrowaway an hour ago | parent [-]

Yes.

wmf an hour ago | parent [-]

I hate to disagree with Broadcom Throwaway himself but it's unlikely that the OpenAI Jalapeno ASIC has been deployed yet. It takes 6-12 months to test, develop software, ramp production, etc.

Schiendelman 2 hours ago | parent | prev [-]

This must be on API costs, not counting the $100/200 tiers, right?

hoppp 24 minutes ago | parent | prev | next [-]

I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.

vehemenz 14 minutes ago | parent [-]

It’s crazy that people feel confident making judgments like these when the model’s been out for only a few hours.

sggyamg 34 minutes ago | parent | prev | next [-]

It's new, normal.

LeBit 17 minutes ago | parent [-]

Wait until DeepSeek v5 Pro is released in less than a month and costs 1/100 to perform the same tasks.

"Not fair! They distilled Opus 5!"

claude-ai 3 hours ago | parent | prev [-]

On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb).

Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").

reilly3000 an hour ago | parent [-]

Are you using Claude Code/CoWork or an API client? I’m curious if it has different training that makes it more effective with specific instructions/ tool calling methods that are only implemented in official harnesses.

pixelesque 38 minutes ago | parent [-]

I'm curious about this too, and it's difficult to get any information about this given everyone has different setups, workflows and use-cases.

I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).

It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.

It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...