| ▲ | JacobAsmuth 4 hours ago | |
Gotta be careful about these benchmarks because they're extremely difficult questions that may be significantly more complex than questions you would ask of the model in real use. If Haiku is "noticing" this and working harder to improve quality, you could still see similar or better cost-per-task in easier domains. ObviousBench is a good test of this. | ||
| ▲ | janalsncm 2 hours ago | parent [-] | |
Agreed. I’m specifically responding to the statement that Haiku was higher on all benchmarks. Normalized by cost (and why wouldn’t you normalize by cost?), it was higher on 1/3. | ||