Your benchmark started being gamed by the frontier models a year ago though. The original idea (find a quirky way to test models with something they don't optimize for) is great, but it needs a refresh.