It is very hard to make sense of these benchmarks given that at least some of these models have probably been trained on the code base of these projects.