| ▲ | epolanski 5 hours ago | ||||||||||||||||||||||||||||||||||||||||
+1, a single test means little. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | jklmnopqrstuvw 5 hours ago | parent [-] | ||||||||||||||||||||||||||||||||||||||||
I don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||