Remix.run Logo
▲ platinumrad 2 hours ago

The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.

▲usef- 25 minutes ago | parent | next [-]

I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.

That would make it about so, I assume?

                     AA     Output  Reasoning  Cost / task
    Kimi K3 Max      44     48k     32k        $2.00
    Half tokens      44     32k     16k        ?

    Opus 5.5 Medium  51     26k     12k        $1.34
    Opus 5.5 High    54     36k     18k        $1.82
    Opus 5.5 Max     58     119k    84k        $5.98

    Sonnet 5.5 Med   41     ?       ?          $0.59
    Sonnet 5.5 High  47     ?       ?          $1.08
    Sonnet 5.5 Max   56     193k    142k       $7.60
 
At Medium effort, it seems like Opus looks more token-efficient while still scoring higher. Medium is usually Anthropic's default setting.

Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem. The increase at Max on Anthropic models is quite considerable though.

▲ 31 minutes ago | parent | prev [-]
[deleted]