| ▲ | hgoel 2 days ago | |
It's tricky to do a proper calculation because everyone obscures what a token or a prompt costs in subscriptions and on top of that we have the complexity of how much different models think. For example, GLM5.3-flash runs slightly slower than Qwen3.8-Next-Flash but uses fewer tokens to think and thus finishes tasks faster overall. On top of that we have complications like Anthropic apparently being extremely misleading about what the 20x plan really means (it isn't 20x the weekly limit of the base plan). I would not be surprised if everyone's doing something manipulative like that. The constant changes to promotional periods, frequent limit resets, harness updates etc make this even more difficult. | ||
| ▲ | esperent a day ago | parent [-] | |
I don't think it's that tricky - there's tools like ccusage that show what you would have paid for Claude at api rates, and I think codex just straight up shows you how many tokens you've used if you run /usage. | ||