| ▲ | Edward40 2 hours ago | |
We initially tried using the mean (average), but a few extreme outliers caused Claude Code's cost to look far higher than it typically is. | ||
| ▲ | joshheitzman an hour ago | parent [-] | |
Claude Code may simply be best used with Anthropic's models and quite bad with Kimi. An alternate solution is to remove Claude Code from the diagram if its so far off from the others that it causes scaling problems. I was looking at the results JSON and it looks like there is only one run of each task with each harness. Since these are disparate tasks using the median means that the headline cost is the cost of one specific task for each harness (or the mean of two it looks like in the case of Exo Harness; I couldn't spot the one task that lined up with the headline cost), but the same task isn't used for the headline cost for each harness. It's not the same as picking a task at random to use as the headline task, but its in the ballpark. | ||