| ▲ | BoorishBears 17 hours ago | |
I've found if you pay close attention, models in Codex get briefly lost on where it is in the plan post-compaction. If it was in the middle of a test it often tries to resume the test from the start, or will seem "surprised" that past steps are completed already. It's hard to believe that sort of confusion doesn't hurt performance a bit, the question is if that degradation is worse than the performance fall off from long-context (which is highly task specific) | ||