That would be very interesting but only the open models allow you to see the reasoning trace
So, open models will be better on this benchmark, which is deserved