| ▲ | daralthus an hour ago | |
omp+0x-alpha beat both cc+fable and codex-sol in creating/refactoring a big eval setup. the former just knows where things should belong and completed the task all the way while the other two failed on both metrics. | ||
| ▲ | dannyw an hour ago | parent [-] | |
to be fair, OMP/Pi is also just a better harness. e.g. https://www.databricks.com/blog/benchmarking-coding-agents-d... | ||