| ▲ | charcircuit 2 hours ago | |
Is no benchmark properly sandboxed? It feels like every single the logs are provided for a benchmark run the LLM is cheating in a way that should have been clearly blocked by a sandbox or the harness. | ||
| ▲ | mpavlov 2 hours ago | parent [-] | |
There's an anecdotal paper 'How We Broke Top AI Agent Benchmarks: And What Comes Next' https://moogician.github.io/blog/2026/trustworthy-benchmarks... | ||