| ▲ | CompoundEyes an hour ago | |
I do think it’s the wizard not the wand at this point given a decent model. These benchmarks don’t have the wizard. Otherwise I wouldn’t see others in the exact same codebase struggle and underutilize agents while others thrive using the exact same ones. | ||
| ▲ | howunfortunate an hour ago | parent [-] | |
In other words, we're still in the era of centaur chess. | ||