There is a theory that the verbosity and comments help getting better results with the current benchmarks. So the models are theoretically getting better but in practice they are getting worse.