| ▲ | esafak 4 hours ago | |
Same task? How did the results compare? | ||
| ▲ | tosh 4 hours ago | parent [-] | |
the tasks were all simple agentic tasks like creating a checksum of a file, merging csvs and so on, fixing a makefile pipeline with known 'good' outcomes all harnesses could reach the outcomes, only cost, time, number of tool uses and so on were different (Claude Code failed once in 1 task but I think that was just an unfortunate outlier, the tasks aren't that difficult) | ||