| ▲ | NitpickLawyer 5 hours ago | |
You really should play the 25 games before stating that it's "simple". The benchmark doesn't just track "completion", it also tracks the number of steps, and the score is based on the median steps took by human players. So in order to get 99% it would mean that the model solved every level of every game in less steps than the median humans. Which, having played the games and having setup harnesses for local models, I find hard to believe. Also the models have to figure out what "end" means. And each game involves some kind of "gotchas" thrown in the harder levels. Some games are only solved by about 2/10 people trying them. The 99% result most likely has some leakage somewhere, either in the preparation of the environments, or from session to session. Seriously, play some of the games. They're fun. | ||