Remix.run Logo
▲ janalsncm 5 hours ago

If we look at their performance per dollar charts,

In OSWorld 2.1 Haiku is better.

On GDPval-AA v2.1 Haiku is equal or worse than Luna.

On Humanity’s Last Exam they don’t seem even be comparing Haiku with Luna.

For these baby distillations of flagships, I expect their users to be very price sensitive.

▲JacobAsmuth 4 hours ago | parent [-]

Gotta be careful about these benchmarks because they're extremely difficult questions that may be significantly more complex than questions you would ask of the model in real use.

If Haiku is "noticing" this and working harder to improve quality, you could still see similar or better cost-per-task in easier domains. ObviousBench is a good test of this.

▲janalsncm 2 hours ago | parent [-]

Agreed. I’m specifically responding to the statement that Haiku was higher on all benchmarks. Normalized by cost (and why wouldn’t you normalize by cost?), it was higher on 1/3.