Remix.run Logo
aarondong 5 hours ago

Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...

midnightbobarun 5 hours ago | parent | next [-]

5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too

impulser_ 2 hours ago | parent | next [-]

It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.

scrlk 2 hours ago | parent | next [-]

Not just compute for OAI, GPT-5.6 is more token efficient across the board vs the Anthropic equivalents: https://artificialanalysis.ai/models?intelligence-index-toke...

No wonder why Tibo can afford to hit the reset button liberally.

charcircuit 2 hours ago | parent | prev [-]

I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

wmf 18 minutes ago | parent [-]

"Price fixing" isn't the correct term here but yes, it's very common to have the same price across different retailers/resellers.

nijave 2 hours ago | parent | prev | next [-]

I think on swebench verified luna was only like 3% points lower for 1/5 the cost

Like 96% vs 93% or something

giancarlostoro 2 hours ago | parent | prev | next [-]

Probably because they made ASICs to run inference for less.

brookst 2 hours ago | parent [-]

Are those actually deployed at scale yet?

brcmthrowaway 2 hours ago | parent [-]

Yes.

wmf 2 hours ago | parent [-]

I hate to disagree with Broadcom Throwaway himself but it's unlikely that the OpenAI Jalapeno ASIC has been deployed yet. It takes 6-12 months to test, develop software, ramp production, etc.

Schiendelman 3 hours ago | parent | prev [-]

This must be on API costs, not counting the $100/200 tiers, right?

anuramat 44 minutes ago | parent [-]

yes; fyi usage limits on the $200 claude sub correspond to at least $1.2k/week in api tokens

eli 2 hours ago | parent | prev [-]

Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.

emmp an hour ago | parent [-]

Indeed, you can filter the graphs to see these the values for alternative reasoning settings of the models. Opus 5 High reasoning scored 59 on the index (exactly the same as GPT 5.6 Sol Max), and costs $1.06 per task (vs $1.04 Sol Max). So these seem essentially equivalent on both metrics.